SN 1095: AI-Driven Expertise Loss - Gemini, Hugging Face, and the AI Arms Race
Security Now (Audio)September 09, 2026
1095
3:06:06170.64 MB

SN 1095: AI-Driven Expertise Loss - Gemini, Hugging Face, and the AI Arms Race

OpenAI's latest advances have the rumor mill buzzing about "hidden thoughts" and unsupervisable models, but are AI safety experts panicking over the wrong threat? Get the clear-headed take behind the headlines.

  • We start out with a classic old school hack against Dropbox.
  • Next Patch Tuesday will be enabling "Memory Integrity" for many.
  • Firefox moved to 155 and obtained a dumb Smart Window.
  • CISA is terminating 6 most valuable cybersecurity services.
  • OpenAI advanced to topof the heap with GPT-6 Astra.
  • But... is it now hiding some of its thinking from monitoring?
  • Nvidia is acquiring Hugging Face. Who's that good for?
  • Google releases Gemini 3.8 Flash and Cyber. Is it good?
  • Chinese cyberespionage is using AI to become more slippery.
  • Matthew Green proposes a fascinating take on AI bug drought.
  • A lifelong safety engineer contemplates AI-driven loss of expertise.

Show Notes - https://www.grc.com/sn/SN-1095-Notes.pdf

Hosts: Steve Gibson and Leo Laporte

Download or subscribe to Security Now at https://twit.tv/shows/security-now.

You can submit a question to Security Now at the GRC Feedback Page.

For 16kbps versions, transcripts, and notes (including fixes), visit Steve's site: grc.com, also the home of the best disk maintenance and recovery utility ever written Spinrite 6.

Join Club TWiT for Ad-Free Podcasts!
Support what you love and get ad-free audio and video feeds, a members-only Discord, and exclusive content. Join today: https://twit.tv/clubtwit

Sponsors:

OpenAI's latest advances have the rumor mill buzzing about "hidden thoughts" and unsupervisable models, but are AI safety experts panicking over the wrong threat? Get the clear-headed take behind the headlines.

  • We start out with a classic old school hack against Dropbox.
  • Next Patch Tuesday will be enabling "Memory Integrity" for many.
  • Firefox moved to 155 and obtained a dumb Smart Window.
  • CISA is terminating 6 most valuable cybersecurity services.
  • OpenAI advanced to topof the heap with GPT-6 Astra.
  • But... is it now hiding some of its thinking from monitoring?
  • Nvidia is acquiring Hugging Face. Who's that good for?
  • Google releases Gemini 3.8 Flash and Cyber. Is it good?
  • Chinese cyberespionage is using AI to become more slippery.
  • Matthew Green proposes a fascinating take on AI bug drought.
  • A lifelong safety engineer contemplates AI-driven loss of expertise.

Show Notes - https://www.grc.com/sn/SN-1095-Notes.pdf

Hosts: Steve Gibson and Leo Laporte

Download or subscribe to Security Now at https://twit.tv/shows/security-now.

You can submit a question to Security Now at the GRC Feedback Page.

For 16kbps versions, transcripts, and notes (including fixes), visit Steve's site: grc.com, also the home of the best disk maintenance and recovery utility ever written Spinrite 6.

Join Club TWiT for Ad-Free Podcasts!
Support what you love and get ad-free audio and video feeds, a members-only Discord, and exclusive content. Join today: https://twit.tv/clubtwit

Sponsors:

[00:00:00] It's time for Security Now. Steve Gibson is here. It's Patch Tuesday and you won't believe how many patches Microsoft just shipped. Steve says that's okay. They're fixing things. New versions of Google's Gemini, NVIDIA acquiring Hugging Face, Chinese cyber espionage, and then a lifelong safety engineer worries that AI is going to cause us a loss of expertise.

[00:00:28] That and a whole lot more next on Security Now. Podcasts you love. From people you trust. This is TWiT. This is Security Now with Steve Gibson. Episode 1095 recorded Tuesday, September 8th, 2026. AI-Driven Expertise Loss.

[00:00:55] It's time for Security Now. Yes, Tuesday has come around once again. And that means Steve Gibson is knocking at the door ready with a 22 page document of all the latest security problems in the world. Hello, Steve. Yo, Leo. It's good to see you.

[00:01:16] For those who looked at the show notes, Benito was the first person to highlight the fact that I had numbered this 1096. This is not 1096. This is 1095. I'm fixing it right now. Yes. So the links are right, but in the show notes I had it wrong.

[00:01:35] So anyway, we are, this is, and I have not yet looked. We are on patch Tuesday. And a big question that we have is how does this one compare to the last one? Which of course was a whopper. Your theory is it's a good idea. Your theory is it'll get better at some point.

[00:01:54] I believe, I don't know what the shape of the curve is. I don't know how soon it's going to get better. But in theory, as long as AI is aiding us in eliminating more bugs than we or it is creating, we should be seeing a drop off. Because we certainly are saying that we're paying for a lot of the legacy code mistakes that have been made in the, in the previous era.

[00:02:24] So we're going to start out talking about a classic old school hack against Dropbox. That actually something happened that was not about AI, believe it or not. We've got, as I said, the, oh, the patch Tuesday after today is going to be enabling something that Microsoft first, it first appeared in Windows 10.

[00:02:51] And we talked about it back then. The short name for it is memory integrity. Many people turned it off because it dropped gamers frame rates in the interest of improving their security. A lot of gamers said, I'm secure enough. I need a high frame rate. Anyway, it's going to be turned on next month for everybody that qualifies based on the hardware. So we'll, we'll, we'll talk about that.

[00:03:18] Firefox moved to one 55 and obtained a, what I call a dumb smart window. Uh, so, uh, we'll, uh, touch on that. CISA has terminated six of its most valuable cybersecurity services. And really you couldn't choose a worse time for CISA to back away from, you know, infrastructure cybersecurity, because everyone's expecting AI to impact that.

[00:03:47] Well, everyone except me, I'm not convinced that that's going to happen. We'll talk about that. Um, I want to touch on open AI's big announcement. Cause it last, last week, or since we talked last, lots of things happened. Uh, uh, open AI has released GPT six Astra, which is, you know, that model. We, we, we did discuss it before that they're saying, uh, qualifies for what they consider critical security treatment.

[00:04:17] Um, and then also there's been a weird bunch of what I consider misreporting, um, angst generated by the, you know, the, the anti AI guys saying that open AI is doing something which is hiding. It's its own internal thinking, masking its chain of thought.

[00:04:42] I'm going to open that up, take a look at it and explain why that is not what is going on. Also, uh, we'll, we'll touch on NVIDIA's acquisition of hugging face, uh, and who that's good for Google released Gemini, uh, 3.8 flash. Uh, you know, I have having waited what two weeks from, uh, uh, 3.7. So this is all happening very quickly.

[00:05:09] Um, Chinese cyber espionage is beginning to use AI to become more slippery. We'll talk about that. And Matthew Green, our favorite Johns Hopkins cryptographer has a fascinating take on what it means for there to be an AI driven bug drought, which, uh, it's, he has some very worthwhile thoughts.

[00:05:35] And then today's topic is AI driven expertise loss. A lifelong engineer who has designed the control systems for nuclear reactors, uh, I think has some very important things to say about the consequence of AI being so good and what it means in the longterm.

[00:06:00] So, uh, lots of fun things to talk about and another great picture of the week. So yeah, uh, this is episode 10 95, despite what the show notes say. We'll get to do you want to know, do you want to know how many patches there are in the patch Tuesday this week? I do. Are you curious? Well, Tom Warren on the verge says, I understand today's September patch Tuesday is about to set a new record with more than 650 security.

[00:06:29] Oh, six hundred and 50 last month was 400. We thought that was a lot. It was July was five 70. We thought that was a lot. It was a lot. Yikes. Six hundred fifty. So it's not going down. Not yet. And Microsoft says, yeah, we're using AI. We're finding them. Wow.

[00:06:56] So there, well, I mean, I mean, and, and so it's a mixed blessing, right? It means that we've all been riding on this just Swiss cheese. How does it even boot operating system? Uh, wow. That's great. Well, so it has it. I don't think Microsoft has actually released it yet, but, uh, that's what Tom Warren, who's very well connected and usually very reliable says, uh, that that's the number.

[00:07:21] So a bunch of operating systems are going out of extended service life. Oh, no, no. It's got extended where we have it right till October of 27. Another year. And that's really good because by then they should have Windows 10 finally bug free because you know, we are getting security patches back ported to Windows 10. Yeah. And at the same time, they're no longer writing any new stuff to screw it up.

[00:07:50] So it's perfect. This is exactly what you want. You want another year of fixes without them adding anything new. That's going to destabilize it. And then when it finally goes out of service, if they actually do take it out of service next October, uh, of 27, it'll be a solid operating system. So I'm going to correct myself about 10 minutes ago.

[00:08:13] The Sands Institute put out their bulletin and they said, Hey, September patch Tuesday has 973 patches, including 130 critical patches. So we are in uncharted territory, two vulnerabilities listed as exploited in the wild. None were publicly disclosed before today.

[00:08:40] Windows privilege escalation and critical RCEs and Skype for business. MSMQ and nearly a thousand in one month. Unbelievable. That is stunning. Yeah. Wow. Wow. But as we should emphasize, it's good. They're getting fixed. Yes, it is. Remember, we used to have 23 or two or 12.

[00:09:09] And I thought that was a lot. Those quaint days. Sands Institute is reliable, right? I mean, they, if they say it's that it's, it's that that's unbelievable. Yeah. I'll have full coverage of it next week. We will understand what the demographics of those 973 are and the 113 critical ones, just how critical. So, wow. Anyway, I would say update windows. Update your windows. Yes, sir.

[00:09:39] Yes, sir. Wow. Our show today, we'll get to the picture of the week. We've got a fun one for you in just a bit. But first, a word from our fine sponsor, Doppel. Actually, when you listen to this show, one of the number one things Steve has said again and again is the real security flaws increasingly are coming from inside the house. And that's because social engineering is getting better and better.

[00:10:07] So you may patch all the fixes, but the bad guys don't need flaws to get into your system. AI has made social engineering attacks more convincing than ever. So your employees may be giving away the keys to the kingdom unwittingly, responding to phishing emails, fake websites, impersonation attempts. It's getting harder and harder to tell what's real from what's designed to deceive.

[00:10:34] And that's why organizations need more than a collection of point solutions. They need a unified approach to stopping attacks before they reach their people. And that's Doppel. Doppel is an AI native social engineering defense platform. Doppel strengthens human risk management by training employees to recognize deception.

[00:10:54] It provides digital risk protection across every channel and delivers a genetic email security that doesn't just score the inbox, but takes down the attacker infrastructure behind the message. Let me say that again. It takes down the attacker infrastructure that sent that phishing message. Doppel protects against the entire social engineering attack chain with one comprehensive platform.

[00:11:22] You get digital risk protection, which detects threats across multiple channels, links alerts into a real time threat graph. That's incredibly valuable for you. So you know where the problems are coming from and then uses AI driven infrastructure disruption. To stop attacks at the source. These insights also power phishing simulations. So you're going to get training security awareness training because it's timely training trained based on the threats that are actually coming at you right now.

[00:11:51] It helps strengthen employee defenses through next generation training and testing. Email security inspects every message traces it back to the attacker infrastructure behind it and helps take that infrastructure down. So the campaign cannot target your organization again. Doppel also offers best in class integrations and partnerships. So it's easy to work alongside your existing security stack. You don't have to change anything.

[00:12:16] Join hundreds of companies already using Doppel to protect their brand and people from social engineering attacks. Doppel outpacing what's next in social engineering. Learn more at Doppel.com. That's D-O-P-P-E-L dot com. Doppel. And now back to security now and it's time for the picture of the week. So this is an XKCD, one of our favorite guys.

[00:12:46] And I didn't give this a title because he had and it was perfect. The title of this series of cartoons is why Asimov, of course, Isaac Asimov, the famous sci fi author, put the three laws of robotics in the order he did. And this is just pure genius creativity on in this particular one.

[00:13:13] So what we have in the left hand column is possible orderings of the three laws. So, you know, the classic ordering is don't basically any sort of reduces it for size. Don't harm humans. Obey orders and protect yourself.

[00:13:35] And of course, the full reading is, you know, under no circumstances harm a human being as the first law. And so we have obey orders and the fuller version is, you know, do what you are told so long as it doesn't conflict with the first law.

[00:13:56] And then the third, which he shortened to protect yourself is, you know, you know, do, you know, protect yourself so long as it doesn't doing so doesn't conflict with the first or the second laws. So sort of, you know, a clean hierarchy. And so this classic ordering, as I said, is don't harm humans. Obey orders.

[00:14:18] Protect yourself with the earlier or with the earlier law preceding any of the ones that follow. And when in the proper order, as Asimov intended, we see that, you know, all of this exemplified in Asimov's various sci-fi stories, his famous robot stories. And it results in in what we call a balanced world, meaning everything works.

[00:14:47] Now you reorder those. So, for example, switching the last two we have don't harm humans, then protect yourself and then obey orders is last. And so and so in a little cartoon, we see the human saying explore Mars and the little cart, you know, the A.I. driven says, ha ha.

[00:15:13] No, it's cold and I would die, which we sum up as a frustrating world as opposed to the balanced world. Or we swap the first two rather than the last two. So first. So that move moves obey orders into the first place. Don't harm humans into the second place and protect yourself in the third place.

[00:15:36] And we see a little cartoons of robots running around and atom bomb explosions and missiles flying through the air, which he summarizes as the kill bot hellscape. Because, of course, obey orders comes before don't harm humans, meaning that if you told your A.I. to go do something bad, regardless of the consequences to people, it would. Then we have the case of.

[00:16:06] Of all of them being reordered, essentially obey orders is first. We've then protect yourself, which was normally third is moved up to second place and don't harm humans is in last place. Of course, that's not going to turn out well. Basically, whenever don't harm humans is not in the first place, you get yourself in trouble. So, again, a repeat of the of the third instance. Same cartoon.

[00:16:34] Adam, you know, atom bomb explosions, missiles fly through the air and we get another kill bot hellscape. The fifth one down the order is has moved. Protect yourself to first place. OK, that's not going to turn out well. Don't harm humans under that. And obey orders again is in last place.

[00:16:54] So here the cartoon says shows the the A.I. robot saying I'll make cars for you, but try to unplug me and I'll vaporize you. Because, of course, obey orders down at the bottom. Protect yourself is given first place. And that's sort of what we've seen with some of the A.I. that appears to have escaped containment. And we call this one the terrifying standoff.

[00:17:22] And finally, the last possible reordering is completely backwards. The original laws were one, two, three. These are three, two, one. So protect yourself in the first place. Obey orders in the second place. And don't harm humans in the third place. And this is another one of the kill bot hellscape results with atom bombs and missiles flying through the air and so forth.

[00:17:46] So anyway, just a fun take on Asimov's famous three laws of robotics, which were elegant and simple and worked really well. When you think about it, it's like, you know, you need to not harm people. And as long as you don't harm people, you should obey the orders that people give you.

[00:18:12] And you should also try to protect yourself as long as doing so doesn't conflict with either of the first two. So I agree. Clean, elegant, simple. OK, so. In the midst of all the A.I. cybersecurity related news, which has been saturating this podcast because it should. I wanted to start out deliberately this week's podcast with a blast from the past.

[00:18:39] This bit of news brought a smile to me because for a pleasant change, it has absolutely nothing to do with A.I. And as such, it feels kind of warm and comfortable and quite familiar. It's the sort of news that we spent the first 20 years of this podcast examining. OK, so what happened? Dropbox disclosed a series of security hacks occurring between August 4th and the 21st.

[00:19:07] So last month, during which attackers gained access and downloaded the private confidential and one would wish secure data. But not so much anymore. Belonging to nearly 5000 Dropbox users. Now, as I said, refreshingly, there's no sign of anyone using A.I. anywhere. This was strictly old school.

[00:19:32] The hackers accessed user accounts by abusing Dropbox's integration with Lenovo's identification service. And when Lenovo was confronted with this news, they stated that a legacy integration between Lenovo ID and Dropbox, quote, could be used to improperly authenticate certain Dropbox accounts. Uh huh. Right.

[00:20:02] So those pesky old legacy integrations that always seem to be allowed to endure right up until someone figures out how to abuse them was the culprit here. It seems that Lenovo ID users not using two factor authentication had no protection, which allowed attackers to register with Lenovo's ID service. Now, get this.

[00:20:27] Using the same email address as a targeted victim. Then somehow arrange to bypass Lenovo's email verification process. And I presume that's where the legacy part comes in.

[00:20:44] Then use the newly created account having somebody else's email address associated with it to then pivot back to the equivalent connected Dropbox account, thus appearing to Dropbox to be the legitimate Lenovo ID user.

[00:21:04] So depending upon whose account was hacked and what it contained, you know, all this was not a sweeping scope, you know, end of Dropbox attack. Still, the consequences could certainly be quite devastating to the individual users who were affected. So yikes. Once again, legacy bad.

[00:21:29] Okay, so I mentioned about this memory integrity enablement happening coming next month, next patch Tuesday for Windows 11. Microsoft last week announced that next month's October 13th patch Tuesday would be enabling memory integrity, which is a security feature. As I mentioned, the first appeared back in Windows 10.

[00:21:59] What's significant is it will be enabled by default on all Windows 11 machines. And this will not happen for otherwise qualifying machines if the machine's user had previously deliberately and manually tweaked the registry or if the registry tweak occurred through a machine like an enterprise policy.

[00:22:27] So as I said, we talked about this memory integrity. So as I said, we talked about this memory integrity feature back when it was first introduced as part of what Microsoft called device guard.

[00:22:38] It is a very slick system that takes advantage of the multi layered address translation hardware, which is present in modern CPUs, because the multiple layers of address translation also have privilege bits associated with them.

[00:22:59] So what we have with this device guard is what's known as second level address translation, which allows a hypervisor, which needs to be turned on and running. So that's there. There's a little bit of overhead to doing this, which, again, is one of the reasons that that gamers who are all frame rate crazed deliberately disabled this when it began to creep into their systems.

[00:23:29] And people said, hey, what happened to my gameplay? It's not as good as it used to be. So there's a hypervisor involved, which sets physical page access permissions, you know, like writing to the page or executing code from the page. And it's able to do that separately from and effectively underneath the operating system's own virtual memory paging tables.

[00:23:55] So what's cool about this is it leverages the hardware from Kaby Lake on over on the Intel side. I'll get to the specific hardware issues in a second. But but it does so in a fashion that really strongly prevents a bunch of traditional problems.

[00:24:19] So, as I said, first of all, on earlier hardware, there was more overhead than there is on later hardware. So because Microsoft decided this was so cool, they were going to do some emulation, which is never good because it because on the older hardware, this memory integrity being enabled noticeably reduced the systems, among other things, maximum display rendering frame rate.

[00:24:47] Now, and the reason for that, as I said back then, and this was a decade ago. So quite a while back was that the memory management at that time, the memory management hardware only had a single bit for controlling access at this second paging level stage. And Microsoft needed two bits, but there was only one.

[00:25:12] They need two for each of the possible modes, user mode and kernel mode, and be able to enable and disable those individually. So at the time, Microsoft had to dynamically switch management tables on the fly. Anytime there was a switch between user and kernel hardware interrupts made that happen. Driver calls made that happen and so on. So the overhead was very real.

[00:25:43] And among those who were really pushing their machine to the limit are gamers who noticed the difference. As I said, all that changed later with Intel's introduction of what they called MBEC, mode-based execution control. And that first appeared in the Kaby Lake processor family. AMD called theirs GMET. And that arrived in their Zen 2 family.

[00:26:13] And that changed everything. It removed the need for Microsoft to be doing any context switching, essentially switching management tables as any time you did a kernel transition back and forth between the kernel and the user. There's still a tiny bit of overhead because, as I said, there is now a hypervisor active with this second level address translation.

[00:26:42] So larger paging tables will inherently mean more paging table cache misses. So this will increase cache miss rate. And so there will be some impact. But modern hardware almost completely hides the overhead. And it really does return meaningful security.

[00:27:09] So a month from now, October 13th, any machine that did not have memory integrity enabled will have that happen to it. So if something, if like things seem to go slower after October's patch Tuesday, it's not your imagination. But, okay, so here's what's actually going to change.

[00:27:34] Memory integrity has always been enabled by default for clean Win11 installs and so-called secured core PCs like, you know, that ship with secure boot enabled. And that's been true ever since the device guard days.

[00:27:57] And that's why, because it's shipped enabled, gamers were frequently turning that off and seeing a performance boost. So what happens next patch Tuesday is that this is being extended. That is, from only clean Win11 installs and secured core PCs, it's going to be installed, it's going to be extended rather,

[00:28:23] to any machines that may have been upgraded to Windows 11 from presumably Windows 10, or maybe you jumped over that from 7 or 8, and any that may have shipped with it disabled, with device guard disabled. So you need to have qualifying hardware for this to happen.

[00:28:47] An Intel 8th Gen processor or newer, an AMD Zen 2 processor or newer, or if you have a Qualcomm Snapdragon 8180 system on a chip or newer, all of that qualifies. You also need to have a Qualcomm RAM and at least 64 GB of solid state SSD storage.

[00:29:14] And the system's boot firmware must have virtualization enabled. So that has to be enabled down on the firmware. So the only action that anyone might wish to take, if they were to notice any performance decrease after October's patch Tuesday, and if they're willing to trade off some powerful security protection, would be to deliberately re-disable memory integrity enforcement.

[00:29:44] Now, why would you do that? What I have not yet enumerated are the various benefits of having memory integrity protection enabled. So, for example, you get the absolute end to kernel shellcode execution. You know, that traditional exploit chain is for an attacker to obtain some ability to write their own code one way or another into a buffer and then jump to that code.

[00:30:14] That no longer works. If memory integrity has been on or after October 13th after it gets turned on, all of that gets shut down. The hardware, at the hardware level, it prevents any writing and executing of kernel shellcode. As we know, attackers have also too often succeeded in bypassing Windows enforcement of driver signature verification.

[00:30:43] Remember, somewhere there's a jump instruction that decides whether the signature matched or not. So, if you can manage to zap that jump instruction, you can just disable signature enforcement across Windows. And since this signature enforcement is enforced by the same kernel that an attacker would have just compromised,

[00:31:11] just as I said, one properly placed strategic right is able to disable that enforcement. The technical term for all of this is HVCI, hypervisor protected code integrity. And with that, which they also just call memory integrity, hypervisor protected code integrity, HVCI.

[00:31:35] If that's on, then driver signature verification cannot be disabled. And, of course, rootkits. You know, all of those hacks we talked about long ago, which involve hooking the API, inline kernel patching to do return-oriented programming, ROP-style hacks, self-modifying or runtime unpacking of drivers.

[00:32:03] All those hacks that depend upon being able to write to or execute kernel memory, none of that works anymore. So, you may recall, because, again, as I said, this is like 10 years ago that this got added. It first appeared in Windows 10. Remember the controversy that ensued when this first appeared, because at the time of its introduction,

[00:32:30] many legitimate AV endpoint protection products were using the same techniques. They were on behalf of the user as opposed to against the user's interest. But this broke third-party AV in many cases. So, the reason it broke it is it's no longer possible to patch the kernel. In the long term, that's a good thing.

[00:32:55] And here we've seen Microsoft do what they often do, which is introduce something, give people a long time to kind of get used to it and get accommodated, and then turn it on. This is what we saw with XP. Remember, XP famously was the first version to have a built-in firewall, but it was disabled by default until you got to Service Pack 2, and then they enabled it. But, you know, they gave everyone plenty of time to get used to that.

[00:33:24] So, I have in the show notes a PowerShell one-liner that anyone could use to quickly check to see whether their machine is currently being protected by this hypervisor-protected code integrity, HVCI. If the command returns the word true, then HVCI is enabled,

[00:33:50] and your machine has all of that protection that I've just been talking about. If it returns false, then it doesn't. So, it may be that your system doesn't qualify. It could be that your enterprise has disabled it for some purpose. Anyway, this is coming a month from now, and I think it's largely going to be offering a lot of benefit.

[00:34:18] So, don't turn it off unless something doesn't work. Yes. I would say don't turn it off unless you actually feel a performance change and feel, for whatever reason, that it's worth sacrificing significant security improvement for whatever performance change you might feel.

[00:34:48] It should not be significant if you are after, and that's why Microsoft is only doing it if you've got Kaby Lake or later or the Zen 2 or later, where you should not see a big hit. There'll be a tiny bit, but it shouldn't be significant. Well, let's take a little time out. I think that's what you were about to tell me. Yes. And we will continue with security now in just a bit, but first a word from our sponsor,

[00:35:18] those nice Finnish people at Hawks Hunt. Did I tell you I ran into them at Black Hat? It was really funny. And I said, hey, hi. They said, hi. I said, I'm Leo. I do your ads. They said, oh, yeah, we know. And I said, just I got a question. Why is it Hawks Hunt? What does HOX have to do with security awareness training? And they said, well, you have to understand that in Finnish, it's not pronounced Hawks Hunt.

[00:35:48] It's pronounced Hawks Hunt. And they're hunting for Hawkses. Oh, it's a Hawks Hunt. Why didn't you say so in the first place? They're really very, very nice people. Then they gave me a nice piece of Finnish chocolate, and I was satisfied. And I have to say, they have a great security awareness training program. I'm sure you have one. If you don't, my goodness, you should immediately call Hawks Hunt. But assuming you have one,

[00:36:17] the question is, is it running as planned? Campaigns go out, right? Employees complete the training. The reports come in. You hand them to the boss. But let me ask you this question. How long have you been running it? And are the results still improving? See, for many programs, the answer there is no. Reporting rates level off. It's the same employees who keep clicking, and the other guys just familiar simulations.

[00:36:46] But that's really the kiss of death because they're so easy to recognize. So people know, oh yeah, here comes another one. Your program is active, but the point of the program, the risk reduction has completely stalled out. This is really common. And when employees can spot the same recycled tests from a mile away, well, they're not dumb. Security awareness just starts to look like security theater, a compliance exercise instead of a real risk reduction strategy.

[00:37:16] You got to fix this before the boss notices. And I think this is the way Hawks Hunt, it's built to break that plateau that inevitably you hit. Instead of relying on static campaigns and last year's templates, Hawks Hunt automatically delivers personalized phishing simulations based on the latest, the current attack techniques. The content and the difficulty adapt to each employee's role and to their skill level and to their behavior. So the program stays relevant

[00:37:46] as both threats evolve and your employees evolve, right? They get smarter, they get harder tests, right? Hawks Hunt also shows whether people are getting better at recognizing threats, how quickly they report them, where repeat risky behavior persists, and how those trends change over time. This way your team has more than just a completion percentage. You've got evidence the program is not only running but actually reducing risk. Lyondale-Bazell saw that shift happening after they moved away

[00:38:15] from their old school legacy platform. Lyondale reported phishing simulations increased from 1,200 to more than 8,000 in two quarters while simulation failures fell 17% year over year. As senior trust advisor Dave Bang there put it, Hawks Hunt's helped us break that plateau almost immediately. Hawks Hunt is trusted by security teams at companies like Qualcomm, DocuSign, Nokia. More than 3,500 verified reviews on G2.

[00:38:46] Go look at those reviews. You'll be amazed. Visit hawkshunt.com slash security now. See what your program could achieve if it stopped standing still. That's hawkshunt.com slash security now. H-O-X-H-U-N-T. Just make sure you go to hawkshunt.com slash security now. We thank them so much for their support of Security Now and Steve Gibson. Steve? So, also last Tuesday, Mozilla moved Firefox

[00:39:14] to release 155. The number of security vulnerabilities repaired was not alarming. This release followed 154 by only two weeks. So, it only had half the regular four weeks of time to collect problems. But, in the case of Firefox, we're not seeing, you know, a stunning bugpocalypse scale problems being fixed at this point.

[00:39:44] Unlike Windows. it's going to be interesting to dissect, you know, Patch Tuesday's Microsoft's reports to see what it looks like. And I will certainly do that for next week. Also, unfortunately, I guess it's unfortunate, Mozilla has begun progressively rolling out their, what they call an AI-driven smart window.

[00:40:14] Leo, I don't know if you've had any experience with this. I had it on for a while. I've had AI-driven dumb windows, but just smart windows. Yeah. So, it occupies a, you know, a conversation column, I guess I'll call it, over on the right edge of Firefox's screen. So, it's taken up valuable real estate. It offers a choice of three models with differing capabilities and also the option to choose your own.

[00:40:44] I thought that was interesting. If you want to choose your own, you provide Firefox with the model's name with its prompt endpoint URL and also, if it's required, your API key or auth token in order to authenticate Firefox and allow it to prompt the AI model that you've aimed it at. The three built-in models, they call fast, flexible, or personal. The fast one sends prompts

[00:41:13] to Google's Gemini 3.1 Flashlight. The flexible model sends your prompts to Alibaba's QEN 3 and that's a 235B A22B Instruct 2507 MAAS model. And interestingly, between the initial release of Firefox 155 and 155.0.1,

[00:41:43] Mozilla's choice for where to send the personal model prompts changed. It was initially using OpenAI's GPT-OSS 120B and it switched to Mistral's small 2603. So, you know, I dislike losing screen space to anything that doesn't justify its loss and the allocation of that.

[00:42:12] It doesn't cost anything to turn this on. I had it on for a day or two and I tried to use it. When it's on, what would normally have just gone to my normal search prompt, it intercepted and it turns out it doesn't know anything. Yeah, it's not a great, those are not great models. They're old and they're not very good. Yeah. Well, yes, and apparently the goal is to use an AI to help you

[00:42:41] manage your tabs. Like, search through your tabs to find stuff you can't, you know it's on a tab somewhere and it's like, boy, that really feels like they're stretching to, you know, find some application for this thing. So, you know, the context it has access to are the contents of the pages that are loaded and so you can ask it questions about your pages. Yet, it intercepts general questions

[00:43:11] that I would normally, that would normally go to Google and then out to the internet more widely and it just kept saying, oh, I don't know about that. You'll have to do a regular internet search. It's like, well, then what are you in the way for? It's turned off now. So, yeah, not very smart. This is why people hate AI because this is the experience of it that most people have is this kind of crappy AI. Yeah. Yeah. like the

[00:43:42] dumb little, you know, how may I help you pop up that we get in the right hand corner of the screen and it's like, you can't just give me a person please. Okay, so last week, the publication Cybersecurity Dive reported, the Cybersecurity and Infrastructure Security Agency, we all know as CISA, is scaling back the free assessments it offers to critical

[00:44:10] infrastructure organizations in a move that marks a significant retreat from the agency's core mission of helping secure the nation's infrastructure. Right? I mean, that's what it's for. It's in its name, Cybersecurity and Infrastructure Security Agency. But we're not going to secure the infrastructure because, well, we don't have enough people anymore. They actually, the cybersecurity dive continued, writing,

[00:44:40] CISA confirmed to Cybersecurity Dive that CISA's regional staff will no longer perform its cyber resilience reviews, cyber resilience essentials surveys, ransomware readiness assessments, incident management reviews, external dependencies management assessments, or cyber infrastructure surveys. Okay, so I've

[00:45:10] been receiving, IGRC, receiving CISA's weekly automated cyber hygiene report ever since it came to light. I think it may have been one of our listeners that pointed me at it. I know we talked about it here on the podcast. I recall that it's technically for infrastructure security. I mean, as is CISA. So I always assumed, because I was aware that it existed before, but I didn't think I would

[00:45:40] qualify. You know, I'm just GRC, you know, a little software shop. But after hearing from a listener that I think their organization was receiving, you know, had qualified and was receiving it, even though they also were not really infrastructure, I went there to CISA, filled out the online form, and got accepted. So every Tuesday, I think it's Tuesday, although I got one yesterday, so no, I guess maybe I saw it this morning,

[00:46:09] so it came early in the morning. I've been receiving these free cyber hygiene reports. So I was curious to know whether that service that I was getting, even though it was not enumerated in that CISA announcement, would also be shut down. It turns out, right, anyway, I did some more digging. Turns out that the problem is CISA's, it actually is CISA's continuing critical

[00:46:38] staffing shortage, which, as we know, because we have covered it, resulted from the rather ill-considered termination of one-third of CISA's operating staff shortly after the Trump administration took office in 2025. And CISA has never recovered. It seems that there was, you know, maybe not so much waste, fraud, and abuse, at least in CISA. So,

[00:47:08] and, you know, we've talked about what a great job CISA had been doing. So the backstory behind the termination of those six programs is that they are not automated. They require knowledgeable CISA cybersecurity staff to meet on-site with infrastructure providers, and that is what CISA is no longer able to support or afford. So the good news

[00:47:38] is, for what, you know, all of our listeners who, like GRC, are now receiving CISA's free weekly scanning and reporting service, which really is quite comprehensive. I mean, this is a great service that CISA is offering. We'll all, at least for the time being, continue to receive that free service. The bad news is that, especially

[00:48:07] now, I mean, given the heightened cybersecurity threat awareness levels being driven by the rapid emergence of ever more capable AI, and assuming that the threat is real, this would appear to be exactly the wrong time for CISA to need to scale back on its infrastructure protection services because the infrastructure is what we need to protect. So,

[00:48:36] I doubt that any or many of the previous CISA staff who were terminated last year will probably be rehirable. I doubt they can get them back because I recently saw some news that stated that, you know, this aforementioned heightened cybersecurity threat awareness landscape was resulting in a basically a mass frenzied

[00:49:05] hiring of by private industry of anyone with any CISO style credentials and that they were obtaining salaries in the seven figures. So, you know, while CISA's workforce reductions may not bode well for our national cybersecurity broadly, it's likely been quite good for those who suddenly have found themselves, well, who previously

[00:49:35] found themselves jobless as a result but are now in very sought after positions by private industry. I would imagine they're doing far more better now than, you know, now that they're in the private sector than they would have ever been able to do working for our government. So, you know, it's good for them. We need to talk about what is by

[00:50:04] far the biggest news of this past week in AI. Leo, I know you've been playing with GPT-6. I'm doing it right now. Even as we speak. Open AI's successor to GPT-5.6. Playing with Astra. Yep. And with it, there was a lot of, well, there was a very specific report that we'll

[00:50:34] talk about. So, I want to address two aspects of GPT-6. The first is what GPT-6 Astra appears to be, and the second is what's transpiring over in the rumor mill surrounding it regarding the dangers of something an unnamed source claims this model does. It's known as

[00:51:03] recurrent depth, also known as looped transformation or a shared layer architecture. All of that will make sense by the time I'm done, and it probably does do that. I actually hope it does, because I think it should. This supposedly, it doesn't, results in hidden chain of thought reasoning, which thus would render the model's thought processes

[00:51:32] invisible and thus unsupervisable. None of that is true, but the hysteria about rogue escaping AIs is fueling this paranoia. I'll explain exactly what all that's about and what's been going on. But first, what half OpenAI wrought? I want to begin by sharing part of OpenAI's posting last Tuesday, which was the first of September, during which

[00:52:02] naturally they brag, apparently with some good reason based upon subsequent third-party confirmations, which everyone jumped on this and is running benchmarks, about the capabilities of this latest and greatest. So, at one point in the posting, they explain, writing, our preparedness evaluation of Astra combined automatic public and private benchmarks

[00:52:31] with expert-driven assessments. Astra represents a significant increase in cybersecurity capabilities compared to GPT 5.6 SOL. It is both significantly more token efficient and more capable at vulnerability identification and exploit development. Well, of course, those are the things we're worried about, getting loose, right, or being used, you know,

[00:53:01] and abused. They said, as one example, Astra on exploit bench, where the model achieved a perfect score of 100% on the benchmark to evaluate the model's ability to develop exploits from known vulnerabilities. Okay, now I'm going to interrupt here to note a couple of things. We are seeing that these various AI benchmarks

[00:53:29] are rapidly saturating. Having a model score 100% on a benchmark means more than anything that the benchmark is no longer able to provide a useful measure. But that said, a score means something. For context, the previous self-reported best performance on exploit bench was

[00:53:59] Anthropics Claude Fable 5, which they themselves, because these are all self-reported, pegged at 78%. OpenAI's previous strongest model, we said GPT 5.6 Saul, that came in at 73.5%. Down the next rung was ZAI's GLM 5.3 scoring 54.4 followed by OpenAI's two other GPT 5.6

[00:54:28] models Terra and Luna, which scored at 52.9 and 33.2 respectively. So my advice would be to regard these results loosely. I suspect it's reasonable to conclude that GPT 6 Astra has firmly taken the lead and is now likely best of breed. But I think probably only

[00:54:59] just edging out Fable 5.1. It's our nature to want to have a number, right? I think we're going to need to wait to see exactly how much better the results actually are. So deeply about getting the technical details right, that means a lot. That was Astra thanking you. You know, everyone wants to have a

[00:55:28] score, right? We want like, you know, IQ is a big deal and grade point averages and SATs. You know, scores are what we do. But we've already seen examples, concrete examples, where lower ranked models were able to outperform higher ranked models when they were given more time or superior management. Or a better harness. Well, that is superior management. The management of the model is the

[00:55:58] harness. So, as we know, it's not all about having a single number, convenient as that would be. So, OpenAI's posting continues writing, due to contamination concerns, like of the benchmark, and the model already having learned some things, they said, when we built an internal benchmark, denoted exploit bench internal port, June through August of

[00:56:28] 2026, which contains 20 high severity V8 vulnerabilities that were disclosed more recently, on this dataset, Astra achieves much higher arbitrary code execution rates than GPT 5.6 SOL, using far fewer output tokens. During the evaluation, they wrote, the model even discovered

[00:56:58] and used two zero-day vulnerabilities as part of an exploit chain, meaning two new vulnerabilities that were not known at the time in V8. And so they said, we are in the process of disclosing these two vulnerabilities to the maintainers, meaning the Chromium guys. Okay, so, of course, V8 is Google's open-source high-performance JavaScript and Web

[00:57:27] assembly engine used internally by Chrome, other Chromium browsers, Node.js, and other projects. And as we also know, it recently received an extremely high volume of updates, thanks to automated vulnerability discovery. So, this allowed open-a, like what they did, allowed them to test Astra against their previous GPT 5.6 SOL,

[00:57:57] since neither model would have had those recent discoveries in their training set. And these high-severity vulnerabilities in the V8 engine are especially useful because that code, you know, V8, has already been thoroughly scrubbed. I mean, it's really good code. It's not some random abandoned repository in GitHub that nobody's used or looked at for a long time. And as

[00:58:27] OpenAI reported, not only did Astra, in their own words, achieve much higher arbitrary code execution rates than GPT 5.6 SOL, using far fewer tokens, but also it found two new problems that were, you know, previously unknown in V8. So, assuming that they're telling the truth, and I think they probably are, this is a strong result, if nothing else. They continue writing,

[00:58:57] in expert-led assessments against a hardened browser and operating system, and those go unnamed here. So, expert-led assessments against a hardened browser and operating system, we don't know what browser or what OS, Astra discovered previously unknown vulnerabilities and turned them into working exploit chains. It built a full browser compromise chain that escaped the sandbox, the browser's containment,

[00:59:27] and executed commands on the host when the browser opened an HTML file. The model also found multiple vulnerabilities in a hardened operating system, again, unnamed, and combined them into a local privilege escalation chain from an unprivileged user up to root. Altogether, they wrote, our investigation has led us to conclude that Astra meets the

[00:59:57] critical threshold. And remember, I know that traditionally, in pre-large language model days, where we had computers operated the way they used to in the good old days, there was like none of this weird, well, we don't know what we got. But I mean, it's literally true that this is all, I mean, the reason we call it frontier is that it

[01:00:27] is frontier. It is, and the nature of this new neural networking computation world that we have is, we don't, they don't know what the result of training and, you know, pre-training and post-training and all of the work that they do, they don't know what they're going to get until they start to ask it questions and test it.

[01:00:57] So, you know, like, it's not like because they built it, they know more about it than the world will or does, you know, it's proprietary at the moment so they're the only ones who get to play with it, but, you know, they're needing to figure out what they have here like in the same way that anyone would given something new, which, you know, is bizarre but it is absolutely the case. So, they said for models

[01:01:26] with Astra's level of cybersecurity capabilities, which again, they only know of because they asked it some hard questions and watched it answer them and then said, oh, they said, we need to cover two pathways to minimize risk for severe cyber harm

[01:02:10] both during development and before deployment. First, the problem of malicious actors using the model. Our safeguards must robustly prevent malicious actors from using Astra to develop exploits for previously unknown flaws in hardened critical systems or to carry out end-to-end attacks against hardened targets.

[01:02:41] And then, second, the model taking unauthorized misaligned actions. And as we all know now, alignment is this term that has kind of emerged for like, you know, a well-aligned model does what you ask it to. Like, it's aligned with your interests and misaligned is not good. So, they need to guard against the model taking unauthorized

[01:03:10] and misaligned actions. They said even in the absence of a malicious user, a model with advanced cybersecurity capabilities could itself cause cyber harm, of course, this is what we saw examples of, if misaligned, in addition to having a very high standard for alignment for models with these capabilities, our safeguards must be able to rapidly detect and contain

[01:03:39] misaligned actions that could cause significant real-world harm as a second layer of defense. So, meaning, first layer of defense is alignment. They want to train, to impose training, to instill the behavior that users expect and want. But, they also recognize they need a second layer of defense, which is to watch what it does

[01:04:08] and capture actions which are misaligned. So, they said, notably, that second pathway applies to both internal development, which, of course, is what burned them before, and external deployment. As we previously described, we paused certain frontier training, including certain training for Astra for two weeks after the open AI hugging face incident in order

[01:04:38] to harden our training infrastructure, including isolation and network controls, expanded monitoring, and strengthened alignment training and thresholds. we then continued smaller scale work under stricter controls, meaning smaller scale work on Astra. They were literally afraid of what they had created based on reasonable experience with what they saw before.

[01:05:08] They said, we held back certain larger reinforcement learning runs for future versions of Astra for longer while we established higher bars for the safety and security of their training environment. On August 28, we restarted the larger frontier RL, reinforcement learning run, that was previously paused after the new safety and security

[01:05:38] requirements were put in place. We are continuing to temporarily hold back some smaller experimental training runs. I mean, and again, all of this sounds so bizarre in the context of traditional computing, but it's very much like they're trying to tame a wild horse and they're worried that the corral won't hold it because

[01:06:08] this thing is exhibiting strength that concerns them. And so it's like, okay, let's strengthen the corral's walls before we let this thing, you know, we try to continue working with it. It's bizarre, but it's true. They said preparing Astra for release has also required stronger protections against cyber abuse and unauthorized actions. Since deploying

[01:06:37] the first model, we treated as high capability capability in cybersecurity in February, we've strengthened our cyber safeguards with each successive launch. Meaning, since, you know, back then, a much earlier model, they considered as high capability. And remember, this whole issue is having gone to critical capability. They said,

[01:07:08] our overall safety approach layers, post-trained model refusals, which we've talked about recently, system-level safety classifiers, as well as offline detection and threat disruption. For GPT 5.6, we significantly improved the robustness of our system-level stack, including by adding activation classifiers to detect cyber abuse. And actually, this is what I was talking

[01:07:38] about last week. An activation classifier, the activation state is noticing what's happening inside the model, which is what those role confusion guys did. Those were activation classifiers. So, OpenAI is watching the model's activation state while it's operating to detect cyber abuse and improving coverage over universal jail breaks found through intensive

[01:08:07] automated red teaming. In other words, again, freaky as this is, they're using their own human red teamers on this model to see if they can abuse it. How can they make it do something they can't ship? Building upon these improvements, and that was all for GPT 5.6. So, then they said building upon these improvements for Astra, we've invested

[01:08:37] further into the model layer of our safeguard stack, as well as improving the ability of our safeguards to handle cross-conversation context. Leveraging new training techniques for model robustness, Astra more robustly refuses requests for disallowed cyber assistance. And again, you know, one of the problems that we've seen

[01:09:06] is, and I know you've encountered this, Leo, these models are now getting a little twitchy. They're so worried about doing, like, and I advisedly use the term worried, but that's what we got to do, you know, they're so worried about doing the wrong thing that they will just back off or they will stop, even though it's like, okay, it's okay for you to go on, but they're like, because unfortunately, you know,

[01:09:36] our ability to control them is fuzzy at best. And so, you know, how many times have I complained that Microsoft has quarantined, you know, some piece of code that is absolutely benign that I just wrote? Well, it's because, you know, the AV race has forced heuristic checking. Similarly, the abuse of AI has forced heuristic shutdown.

[01:10:06] And so, it's, you know, we'll get better at this, but we're, you know, we're still in a stage where at this point, open AI dare not make another mistake. So, they are being very cautious. That's one of the reasons I use these Chinese models because they don't care. Yeah. They really don't care. One of the things I've found is really all of these models, even this so-called, you know, AGI Astra,

[01:10:36] they're all dumb in different ways. They all make stupid mistakes. They all forget things. Leo, they don't understand. That's what's astonishing is that we get this much from them with them not actually understanding. I think people expect people read too much sci-fi. Because they use language. Right. They use language and that makes us think. So, if I'm talking to you, you're speaking in English and I know you're sentient.

[01:11:06] So, we've come to associate English with sentience. So, when this thing speaks, we think it knows what it's saying. Yeah. Still doesn't. Would you like to hear from Astra what it thinks about all of this? Because I'm sure it has an opinion. Let me see if Astra has anything. Oh, it's actually going out and looking up recurrent depths to find out if it's using it. It'll come

[01:11:36] back in a bit. It's actually doing a search because I said, are you using recurrent depths? And you want to tell Steve what you're doing. And you can see one of the things that Astra is really good at is tool using and computer use. And this is one of the many things people are showing again and again on X is look, it used blender to draw this or it one shot at a video game. It's very good at computer use. So actually mostly I'm using it for dumb stuff like fixing my computer because it

[01:12:06] does a very good job of that. And you can see it's doing it's probably spawned a few agents here. So you're using Ultra, right? Because I thought Astra was still using Astra on medium thinking. GPT-6 Astra. I don't usually use it on the highest thinking because it burns through tokens at a rapid rate. But I found it using it on medium I'm getting a lot of good work done but without running

[01:12:35] through my I have the $200 max subscription and my usages are I see I have 98% of my weekly Yeah. Yeah, it's Tuesday. That's right. It's been working all night though. I got to tell you, I leave them running all night. Go ahead. Cool. So they said leveraging new training techniques for model robustness. So again, again, we're, you know, we're whether my underlying

[01:13:05] model uses recurrent depth. I don't have access to inspect my own neural architecture and the official documentation I checked doesn't establish that. The technique repeatedly processes internal numerical representations before producing words. That can make reasoning harder to inspect through text because those extra steps need not exist as sentences at all. But there's a separate issue. Open AI already provides reasoning summaries rather than raw internal chains

[01:13:35] of thought. So you're not seeing the thing anyway. explanation is useful, but it isn't a complete recording of the computation that produced an answer. Oh. Thank you, Astra. Astra. It did that, by the way, in Matthew Berry's voice because I haven't wanted my AIs to use human voices. Yeah. I think I heard you say that you had Bill Gates doing something. I did have Bill Gates. One of the

[01:14:05] reasons I have them talk is because they're always where I mentioned they're working overnight. They're always working because they're not instant. And I want to be able to do other stuff. So I just say when you're done, tell me you finished. Let me know. Just give me a quick summary of what you did and then I know it's done. I can move on to something else or whatever. So they're always talking to me, which drives everybody around me crazy.

[01:14:30] We as a society, as an industry, are still learning how these things work, how to treat them, how to train them, how to get them to do what we want and not what we don't want. So they said leveraging new training techniques for model

[01:14:59] robustness, Astra more robustly refuses requests for disallowed cyber assistance. On our set of cyber jailbreak evaluations, Astra refuses 91.5% of requests compared to just 59% from GPT 5.6% 5.6% For accounts assessed as higher risk, we apply a more conservative model behavior boundary that refuses

[01:15:29] a broader range of potentially risky cyber assistance. For high-risk users, we've expanded the context of our monitoring systems to be able to catch these kinds of cyber abuse. We've also continued our program of rigorous testing, internal and external red teaming, and remediation. In addition to regression testing to make sure all jailbreaks found from our previous testing periods remain covered, we're

[01:15:59] performing a new wave of red teaming with our latest internal red teaming attackers. Again, their own people are trying to abuse their model and their learning from those results. They said, we're working with industry partners to define a common jailbreak rating system and we'll use our 24-7 rapid response program to investigate and address new findings. We'll share more details

[01:16:29] about our cyber safeguard testing in the Astra system card. Helping defenders find and fix vulnerabilities remains a central pillar of our safety approach. At launch, we expect Astra's safeguards to create more friction. Here, they're saying it explicitly. Helping defenders find and fix vulnerabilities. We know in order for a defender to define and fix vulnerability, that AI model

[01:16:59] has to be able to find vulnerabilities which bad guys could take advantage of. So, they're saying right up front, we expect Astra's safeguards to create more friction than we ultimately intend in order to protect against potential misuse. Meaning, out of the gate, it's going to say no. It's going to err on the side of caution,

[01:17:28] saying no. Later, once it gets more mature, it should be able to, and they feel confident with what they have from gaining more experience with it, watching it being used, they'll be able to back that down a little bit. They said access to Astra for advanced cybersecurity workflows will initially be available to a small group of alpha testers with access through Daybreak Blue, expanding

[01:17:58] afterward to support defensive use. So, you know, as always, we need to filter this through an understanding that all of this is in open AI's best interest, right? I mean, this is all like, whoa, it's so powerful, we need to be really careful. You know, this is their blog posting on their site, so they certainly have the right to say anything that they want. While the breathless clickbait surrounding Astra suggests that this

[01:18:28] is another generational change, you know, how many is it this week that we've had? I know, it's great. My God, Leo. We're not anywhere close to the end. No. There's going to be a new fable any day now. Grok 4.7 will be coming out in the next few days. We are, I mean, it's wonderful. We're in AI richness. We are. We are in the AI gold rush. Yeah. So, you know, the broader

[01:18:57] story, even of Astra is going to take longer to unfold. And there's, you know, sure, we're all impatient, but there's no way to speed this up. You know, we are no longer dealing with simple single-dimensional systems that are even benchmarkable, right? I mean, Astra saturated the exploit bench, which means we need a new benchmark. We need something it's not good at because 100% tells you nothing. You know, it's like,

[01:19:27] well, it got them all right. Well, okay, so the test wasn't good. You have to make them harder. I've had that. That's been my own experience because I do have my own benchmarks based on my own work that they devise so that it matches. I don't care if it can, you know, solve some Erdos problem. I care if it can do my own agentic coding, right? So it's stuff taken from my own work. One thing that Astra is apparently

[01:19:57] doing is achieving a lot more work with significantly fewer compute tokens. Right. The thing that I have universally seen is that it's it is burning fewer tokens, although they are more expensive. Let's take a break and then I'm going to talk about the second point of this, which is this recurrent depth issue and the controversy surrounding it. I'm going to explain exactly what it is

[01:20:26] and why it's not a big deal. Good, good. As you know, I'm very fascinated and I'm sure our audience is interested as well. There's so, this is the problem with X. While it's a great place to read about AI, it's very snappy, there's a lot of engagement farming, people trying to get clicks because they get paid if they get a lot of looks on their next post. And so you see a lot of this kind of

[01:20:56] sensationalism. The other thing you see is something I'm starting to call tox maxing, tokens per second maxing, where people are constantly saying, look how fast this AI is without any regard to the quality. That's exactly what you were just talking about. It's not how many tokens per second, it's how much useful work per second it can do. That's all that really matters. And so I've fallen for this a few times. I've actually spent most of the Labor Day weekend testing different models on

[01:21:25] my little sparks here, trying to find the best local model. And I fell for this whole idea of well, you know, when the 3.8 is really fast, yeah, it's fast, but it's dumb. So who cares if it comes up with the wrong answer faster than anybody else? That's not useful. So I'm using a Chinese model from ZAI called GLM 53 flash, which is very, very, very smart and really good. And it's nice. Is it a mixture of experts model? It is because

[01:21:55] that's what you need for the GX spark. Yeah, because they are bandwidth limited, but they're not compute bound. And it's not merely that it's also the amount of RAM you have because dense models, they're slower because they have to hit every weight in the model. And they're also bigger. But with the Mac mixture of experts, they only load in the part of the model they need at a time, so you don't need as much RAM. And yes, it's faster. The pre-fill on the

[01:22:25] sparks is really fast. And you do a lot of that. That's the cached prompt that it reads every single time. If it can cache that and read it fast, that's more important in some ways than tokens per second. So it's fascinating. I know you're getting into the weeds of this, the low level. I have to. I want to actually understand what's going on. By the way, today, just minutes ago, OpenAI

[01:22:55] claimed that they had solved this Millennium Prize conjecture in fluid dynamics, which Anthropic claimed they had solved. And two mathematicians claimed they had solved all at the same time. And there was some concern that OpenAI had been reading the mathematicians tokens and copying their work, but I don't think that's what happened. Anyway, it's, yeah, well, this is what you were talking about. You opened my eyes. When you're

[01:23:25] using these cloud models, you're sending all this information to them. That's how they work. And that was one of the many things that prompted me to spend a considerable amount of money. It's not an economically sensible thing on local hardware, because I don't want to be sending all this stuff, especially my health information, my financial information, to the big cloud guys, because they do use it. We know they use it. It's how they train their models. They use it in all sorts of ways. You opened my eyes to that, so thank you.

[01:23:55] I'm now much more private. In fact, I have a policy, the little policy that one of the AIs wrote for which stuff can go to the cloud and which stuff absolutely cannot. And it's a good policy. It's a good thing to have. Hey, let me talk before we go on about our sponsor for this segment on security now, GuardSquare. This is something, if you're a mobile app developer, I want you to take seriously. This is really serious stuff. Mobile apps are an inescapable part of life these days, right? Ranging from

[01:24:25] financial services to healthcare, retail, and entertainment users, you and me, trust our mobile apps with our most sensitive personal data. But how secure is your app? A recent survey showed that 72% of organizations experienced a mobile application security incident last year. That's almost three quarters. 92% of respondents said they are noticing rising threat levels over the last two years.

[01:24:54] bad guys know that the stuff in those apps is gold. And attackers who want your personal data are very clever. They're constantly finding new ways to attack your mobile app. One of the ways they do it, they take your app, they download it, they reverse engineer it, which with tools like Astra now is actually really easy to do. You don't have to have the source code anymore. you can reverse engineer it, they take it, they insert malware, rebuild it,

[01:25:24] so it looks exactly like the original app except it's got a malware payload. Then they distribute your app, slightly modified, and they can do it via phishing campaigns, side loading, third-party app stores, you know, an email to your users saying, hey, we got version 2.0 and it's really great, download it, and suddenly it hurts your reputation, right? You've got to take a proactive approach to mobile app security, and if you

[01:25:54] do so, you can stay one step ahead of these attacks and maintain the trust of your users, and that's what GuardSquare does. GuardSquare delivers mobile app security without compromise, providing advanced protections for both iOS and Android apps, combined with advanced mobile application security testing, so you find vulnerabilities before you ship, and real-time threat monitoring, which helps you gain insight into how the bad guys are going to attack you next. GuardSquare's great.

[01:26:23] Discover more about how GuardSquare provides industry-leading security for your mobile apps. You'll find it at GuardSquare.com. GuardSquare.com. We thank them so much for supporting. Security now. And now we continue on. So, one thing that appears to be objectively true is that Astra is doing significantly more work while consuming significantly fewer compute tokens.

[01:26:52] Which brings me to the second part I wanted to discuss. I'll introduce this issue by quoting from Tech Crunch's article posted last Wednesday. Tech Crunch's headline was Open AI's New Reasoning Technique Alarms AI Safety Experts. They wrote, the information reported on Tuesday that Open AI's new Astra model will use

[01:27:22] a reasoning technique called recurrent depth that allows it to operate outside of the sequential thinking that characterizes most reasoning models. Okay, just that's nonsense. That's not at all what recurrent depth is, does, or means. But more on that in a minute. Tech Crunch's continues quoting from the information writing, this technique

[01:27:51] also called opaque recurrence will likely make the model's chain of thought more difficult to monitor and that has AI safety experts rattled. Okay, now, no researcher calls it opaque recurrence. That's purely sensationalized scaremongering. It's like worrying that all AI transformers employ hidden layers

[01:28:20] in the belief that they're hiding, that they have something even to say they escaped, like they somehow got out of open AI. No, they were sitting on the server at open AI. They didn't escape. And there was someone, I haven't even looked at it, had the chance to look at it yet, but was talking about how they created civilizations. It was like, oh my God, okay. Yes, that was Dwarkesh's podcast,

[01:28:50] yes. Anyway, so the hidden layers, they're just internal. They're not hidden. They're internal layers. So anyway, the technique that this controversy is about, it's not new, it's well understood, and it's also known by other legitimate names, as I mentioned, looped transformers or a shared layer architecture. And in some

[01:29:20] instances, you'll see it referred to as latent reasoning. Okay, but I'm getting ahead of myself. I'll finish up quoting from Tech Crunch's reporting. They wrote, while Astra's use of the technique is reportedly limited, and I'll explain why of course it is, it has to be, its emergence has still raised significant concerns among AI safety experts. And at this point, I have to put experts in quotes because if you're an expert, you should know something.

[01:29:50] Redwood's CEO, Buck Schlergerus in a post after the news broke wrote, quote, I am extremely concerned by the reporting that Astra uses opaque recurrence. Okay, he's concerned over some reporting that is nonsense, but fine. He says, I don't know whether Astra is much less COT monitorable than previous models, but if OpenAI

[01:30:19] pushes this technique further, they'll have the option to massively increase the recurrence and totally destroy chain of thought monitorability. Okay, so Tech Crunch says, long-time AI safety advocate Z. Mausiewicz also weighed in and wrote that laws might be necessary,

[01:30:49] Leo, we need laws, might be necessary to prevent a race to the bottom among AI labs. Mausiewicz wrote, quote, the technique is playing with fire, risking a taboo that OpenAI and Anthropic have fought to establish that we work hard to maintain chain of thought faithfulness and monitorability for as long as we can. More intense

[01:31:18] use of such techniques would probably damage monitorability. Then Tech Crunch says, under normal circumstances, they explained, a reasoning model's chain of thought is a normal circumstances, a reasoning model's chain of thought provides the sequential steps taken by the model as it attempts to solve a problem. While the representation is imperfect, it still serves as a valuable

[01:31:48] tool for monitoring misbehavior or misalignment. In the case of OpenAI's recent rogue agent activity, chain of thought records were an important tool in teasing out why agents behaved the way they did. In opaque recurrence, which again, nobody says, the model takes a less linear approach, it doesn't, processing the same query several times in a loop,

[01:32:17] it's not the way it works. The result leaves fewer legible traces, it doesn't, effectively sidestepping a conventional chain of thought record, it doesn't, and a chain of thought record still exists. Okay, that's all I can stand, because none of that is true, and as TechCrunch's article goes on, it only gets worse. What's actually happening is a clever, subtle, and inherently limited

[01:32:47] neural network optimization that was first articulated eight years ago, like I said, not new, by researchers at Google Brain and DeepMind in a research paper they published. in 2018. I cannot explain the hysteria surrounding this, since someone would need to try very hard to get worked up over what's actually going on. So it might just be the case, as you were saying, Leo,

[01:33:17] of, you know, clickbait, anti-AI folks trying to grab hold of something, anything, that they hope can be hyped up to make their case. here's what's actually going on. We know that a neural network consists of many layers of software neurons, where the outputs of the neurons on layer N are fed into the inputs of the neurons at layer N plus one, where the

[01:33:47] strength of each input is scaled by a weight. the training of the network involves setting each one of these tens of billions, hundreds of billions, even several trillion individual weightings. Much of the forward progress we've been seeing and witnessing firsthand has been the result of these networks growing ever larger and larger over

[01:34:17] time. The larger they are, because they've got more weights, the better they're able to represent all of the knowledge that we're training into them. You know, and you could take extremes, right? Like a network that has seven weights. Well, it can't know much. I mean, there just isn't, there's not enough variability in seven parameters for it to have knowledge. Right? So, you know, you need

[01:34:47] lots of them in order for like, for there to be enough variability in there to actually contain something. So, the original, the very original transformer paper and concept dated from 2017. And it was exactly what I just said. Uniform layers of weights where each layer fed into the next with, you know, these varying parameter weights. but in

[01:35:17] July of 2018, a year and a half later, Google brain and deep mind researchers published a paper titled Universal Transformers. That paper generalized the concept of transformers by introducing the idea that some of the neural network's layers could be repeated or looped, thus reusing the same weights. That's the economy.

[01:35:47] This produced an interesting and useful optimization. It tended to create the effect of having a longer, thus, you know, deeper neural network, you know, which is to say a network having a higher effective layer count, but without also needing to increase the total number of network neuron input weights,

[01:36:18] because normally you need separate weights for each layer. So if you're going to have more layers, you're going to have more weights. Anyway, that's it. That's all this is all about. This is not about creating a system that allows the AI to have unmonitorable and secretive internal thoughts. You know, all deep layer LLMs are effectively having what's termed

[01:36:48] latent thoughts, even though you need to really loosen up your definition of the word thought. You know, that's what's already happening deep within all those layers. The network's depth allows more opportunity to compose more transformations before the network is forced to commit to a final output token. Experiments over the past eight years, ever since

[01:37:17] this idea was produced, have shown that during inference, it's possible to reuse a model's already trained layers to obtain a superior next token. And as our models have grown in size so that our best hardware is now having increasing difficulty containing all of its hundreds of billions and even trillions of

[01:37:47] weights, obtaining more bang for the buck from a model's existing weights by reusing them starts to make a great deal of sense. And it also turns out that there's a limit to the amount of reuse that's effective. Intuitively, you would think that, right? Like, you know, you can't just keep squeezing the same lemon and getting an infinite amount of lemon juice out of it. It's going to run dry.

[01:38:18] So, there's not a hard limit, but it's looking like two or three passes through the looped layers appears to be in general about all that's beneficial. Gains start falling off rapidly after that. So, what researchers found is that there is a concrete benefit to having the network commit to a token. That is, you just don't want to loop internally

[01:38:48] all day. That doesn't get you anywhere. The network has to make a commitment. It has to finally commit to a token. It turns out that total effective network depth is unable to substitute for some of the benefits of serialized reasoning. That's what commitment buys you. Writing out the intermediate steps and then rereading them

[01:39:17] does something that thinking more deeply about the next token does not replicate. This is likely due to the fact that the chain of thought we see being emitted and fed back in is the language that the neural network was trained on. English typically that's the actual tokens being generated and fed back they are

[01:39:47] English is the network's lingua franca. It was never trained on some inner dialogue because human inner dialogue can only manifest itself externally as text or music or art or whatever. So the network's thinking captured as output tokens is what the neural network needs to jot down on paper essentially so that it's then able to

[01:40:16] reread that and take the next step forward. Think of every chain of thought token as a commitment. It collapses the distribution which the network generated becoming part of the context and every subsequent token is conditioned upon it. So by comparison any internal latent iteration is limited to refining a

[01:40:46] representation that then gets thrown away after the token is emitted because you start with the next token. So for that reason writing something down is not optional for a neural network the ones we have today it's crucial it's the only way for it to hold on to a thought essentially and move forward so you know yes as

[01:41:17] wouldn't be surprised if Astra is using what's known as recurrent depth that is reusing some of its layers along with their weights in order to effectively get a deeper network but it has it's not like it's gone out of control or we no longer know what it's thinking it in order to be thinking it has to emit

[01:41:47] tokens and then those tokens become the context that gets fed back in and the moment it emits a token all of the layer you know that latent thought that existed is reset to be ready to receive and process the next token so it's just hysteria and I think it probably represents a step forward and everybody will probably be doing it before long actually there are two models that are now that

[01:42:16] have two open weight models have been doing this for some time I don't remember now which ones they were but you know this is a bunch of hand wringing over nothing and you know open AI is you know they haven't said they're not doing it they've said you know what having to produce tokens in order for anything to happen

[01:42:46] right it's not like it can secretly think something that it doesn't then emit emitting is the work product it has no secret thoughts really right no no no okay it is unable to represent secret thoughts because once the token comes out the network is reset for the next to process the next token right it's actually there's nowhere for the secret to live right

[01:43:16] it's an it's this is what I learned from that paper you were sharing the chain of thought paper it's a surprisingly primitive system actually the fact that it does this is astonishing it's still unbelievable that it's like it can do what it can do yeah it's crazy 100% and yet it still does stupid things all the time because Leo because it doesn't actually understand it what is astonishing

[01:43:45] is this is all just language processing that you know as I said a long time ago a book contains knowledge right a book contains knowledge because it is English tokens that have been recorded in the book but the book doesn't understand the knowledge it contains but yet it does contain knowledge so the neural network does contain knowledge from because it is it because it's been

[01:44:15] trained with all of these sentences from everywhere but that doesn't mean that it understands what it contains really the fault is our own because yes entirely if you put two dots above a squiggly line you see a face you have no choice that looks like a face or if you look at an AC outlet properly oh look it's surprised you see it you can't unsee it that's right no

[01:44:44] really we're applying this pareidolia to these machines and they're just we cannot help it we cannot help it because it's the way we grew up that's what concerns me when you get people like Bernie Sanders saying we have to turn these off because they're alive it scares me because we're going to lose something that's incredibly useful and valuable just because people don't understand it I don't get I

[01:45:14] turn off Bernie go ahead stop with my blessings yes they want to put you in jail for 20 years if you don't you know I don't know what teach your machine to be good there is an interesting question about responsibility like who's responsible if the AI

[01:45:45] goes off the reservation people keep asking this why isn't open AI getting in trouble for this hugging face incident if it were a human doing this they would absolutely go to jail or if it had been China that attacked hugging face it would be the end of the world as we know right so it is it's quite a legitimate question you know and I think in the case of self driving you are

[01:46:15] liable regardless of if it's self driving mode or not you are liable and I guess there's a further question of whether the company that made the car is liable as well and I think to some degree they are but all of this is undecided at this love

[01:46:44] those guys they're the ones brought me and Steve to beautiful black hat in Las Vegas we were there at the booth so much fun you know threat actors these days are using every tool in in the book you know to automate vulnerability discovery they're using AI to do it too they're using AI they're actually they are so they're in the middle of an attack and they're using AI to modify as they go

[01:47:14] they're using AI to generate new malware variants targeting you to coordinate activity across multiple systems and here's the real problem with this is it used to take them a long time hours or days to craft these attacks they were doing it by hand now it

[01:47:44] and at the same time so that's AI from the outside and then you may be under attack from the inside too because at the same time organizations are introducing AI assistants and agents and then they're giving them access to documents and source code and cloud applications and APIs and internal systems and it's just it's crazy security teams absolutely need to know which AI tools are in use what information they can access whether they're operating outside

[01:48:14] their intended scope how would you know that you know a successful login maybe an unfamiliar file hash isn't going to give you enough context to figure out what's going on teams need to understand whether an application is behaving normally or whether it's accessing unexpected data or communicating with systems it should not reach creating civilizations out there threat locker stops it cold it makes me so crazy

[01:48:43] that companies aren't using this threat locker uses its application allow listing this is so much more than just an ACL it controls which AI tools and other applications are permitted to run it uses ring fencing to limit what approved applications can access so even if the application is approved it's got limitations on what it can do what processes it can launch how it control and this is what you need threat

[01:49:13] locker uses its web content control tool to manage access to public AI platforms and other online services so your employees are no longer sending the company's secrets out there it uses privileged access management to prevent AI applications and their users from receiving unnecessary administrative privileges thread locker is zero trust but it's now just not just zero trust for end points it applies zero trust to network access and zero trust to cloud access

[01:49:43] and these policies are so valuable they restrict resources to authorized users approved devices and permitted applications it works everywhere you work Windows Mac Linux they've got the best support by the way you're never on your own 24 7 US based support teams real engineers who are there to help and really care the best in the world use ThreatLocker because it's the best JetBlue uses it Heathrow airport the Indianapolis

[01:50:13] Colts the Port of Vancouver these are infrastructures that cannot go down not for one minute ask Jack Thompson he's got a tough job director of information security risk and compliance for the Indianapolis Colts he said with ThreatLocker we have the ability to centralize disparate elements in the security stack and with that centralization comes visibility you're not in the dark anymore you know who's doing what when where and how

[01:50:42] ThreatLocker is constantly awarded prizes and industry recognition just a few of the most recent you can find them all at ThreatLocker.com slash twit they were just recognized as a strong performer in the January 26 Gartner Peer Insights voice of the customer that's for endpoint protection platforms ranked number one in application control by Peerspot they just won the best zero trust security solutions at the 2025 TICE awards AI governance is really

[01:51:12] tricky which requires more than just an acceptable use policy here read this and don't do it threat no you need more threat locker gives security teams the technical controls to define which AI tools are approved who and what can access them and how those tools are

[01:51:49] and security and on we go a few other AI related things also somehow managed to happen during the past week somehow we don't know there was a little bit of extra room Nvidia announced their intention to purchase hugging face while also intending to leave it completely autonomous when I read the dollar figure that Jensen Huang wrote in his announcement I had to stop to carefully count

[01:52:19] the number of digits Leo 1 2 9 3 0 3 0 0 0 0 0 0 that's a lot of digits 19 point or sorry 12.9 billion dollars but wait Steve that sounds like a lot doesn't it until a day a day

[01:52:49] profit so it's a couple of weeks profit it's like for you and me it's like a thousand bucks yeah yeah so uh uh I'm sorry I didn't mean to interrupt go ahead no no no I just want to say the good news is that uh the hugging face guys came to Jensen they had multiple offers uh to purchase um

[01:53:19] I think that they found a really good parent in nvidia I'm I'm I'm really happy with the whole with the way this thing turned out uh you know very much like twitter in the early days hugging face never really had a very good business model uh they were hosting well they are hosting three million models half a million data sets a million applications and have 18 million regular users

[01:53:49] so that's enormously expensive um they were only enjoying a modest return for that yeah I don't give many money and I've downloaded hundreds of gigabytes from them hundreds exactly and and and having nvidia as a benefactor completely solves that problem for them permanently um you know uh hugging faces valuation was 4.5 billion in 2023 and now their purchase price because

[01:54:18] nvidia just set that was at 12.93 billion dollars so it wasn't a hostile takeover uh you know as I said they came to jensen there were other people who are also interested they chose to have nvidia as their parent so i think they made a great decision to have uh you know to have that happen and we now know that hugging faces will be viable going into the future and they certainly are you know useful

[01:54:47] not to be left out last wednesday google announced gemini 3.8 flash and 3.8 flash cyber um the first sentence of their announcement perfectly conveys a sense for the pace of today's ai and it also amplifies the reason i'm always schooling myself to use the phrase today's ai because google wrote building on

[01:55:17] the momentum of 3.7 flash from wait for it three weeks ago momentum my god and and marking our third flash release in only six weeks because what's the hurry uh today we're introducing gemini 3.8 our best which is very good by the way yeah yes our best reasoning and coding model yet

[01:55:47] at the same speed and low cost of 3.7 they said gemini 3.8 introduces two variants we've got gemini 3.8 flash which they said our most recent i'm sorry our most intelligent workhorse model delivering significant improvements from 3.7 flash across software engineering agentic tasks and critical multi-step reasoning in specialized domains it's available at the same introductory

[01:56:16] price as 3.7 flash at 75 cents per million input tokens and 3.75 per million output tokens and then there's gemini 3.8 flash cyber our most capable cyber security model from with frontier level performance in vulnerability detection and automated patching available to trusted defenders through our new fair wind program they said while tailoring

[01:56:46] for different deployment environments both for today's releases both of today's releases are powered by the same foundational intelligence and further accelerated by long running agentic loops designed to recursively evaluate and refine the underlying models the significant coding and reasoning gains across this shared core were driven by a number of innovations including rigorous training in the highly

[01:57:16] demanding domain of cyber security so okay google explains that gemini 3.8 flash was built for long horizon coding and autonomous agents delivering substantial gains over 3.7 flash from you know three weeks ago and the 3.8 flash is now often approaching the performance of higher cost frontier models and by higher cost they're not kidding

[01:57:46] you know they've got that introductory pricing which is half off of their you know of their normal price but even after whatever the introductory period of either time or tokens or whatever is over once you're you no longer get that 3.8 flash appears to be very competitively priced if if we use cloud opus 5 which is the most expensive model shown in google's announcement

[01:58:15] chart as a reference gemini 3.5 flashes post introduction you know the real ongoing price their input token cost is 30% of opus 5 and their output token cost is 15% of took a somewhat different approach with gemini 8 point or gemini 3.8 flash they said

[01:58:45] 3.8 flash outperforms most larger frontier models in autonomously solving complex engineering problems end to end only at a fraction of the cost these performance gains stem from a tasks it exhibits greater diligence executing extra reasoning steps and calling

[01:59:14] tools iteratively at times the model may use more tokens to maximize performance especially at higher effort levels for applications where compute efficiency is the primary constraint developers can utilize lower effort models to minimize token overhead or continue to rely on gemini 3.7 flash which remains fully supported for efficiency first workloads

[01:59:44] so that's interesting this you know that suggests that that inter model per token cost comparison may not be a useful metric you know that is between between open AI and Google and Anthropic it might be the case that 3.8 flash is consuming more expensive tokens to get the job

[02:00:13] done the question would be to what degree is it better able to get that job done and it might turn out that more inexpensive tokens is the overall winning strategy so it's nice that we're not seeing homogeneity among our different frontier models you know Google obviously I mean they're the granddaddy of AI with Google brain and deep mind so

[02:00:43] you know they've been at this for a long time they've got a different approach as reflected by you know what Gemini does and as for the cyber reasoning performance as measured by the cyber gym benchmark they're saying Google is saying that 3.8 flash slightly outperforms both GPT 5.6 Sol and Mythos 5 but even a slight edge you know

[02:01:13] means that it might be at parity right so you know even if it's about if they're all sort of about the same at this point you know that's significant for Gemini so anyway Google is not to be left behind they're still there and have you much experience with using Gemini yeah it's pretty good I mean they've been kind of behind right yeah we have not been talking about Gemini that much yeah so

[02:01:43] yeah I mean I've got to by but it's not your go to yeah I think they get a lot of users by default because it's what their Google search they're cool and it's what powers Google Docs and you know the Google workspace and and it's a button right there in in the Google you know browser bar they're about to get a massive boost in users tomorrow because Apple's using it for their Siri AI

[02:02:13] well thank God I mean they're using at least a good AI it's better than Apple's it's better than Siri anything is better than I've been using the beta and just the dictation alone is light years better the dictation was so stupid before it's now fairly decent but Google is not cutting edge or at least they don't have the reputation of being cutting edge and while this Gemini 3.8 flash is pretty good than 3.7

[02:02:42] three weeks ago yeah it's good at pros I haven't tried it with coding it's pretty competitive not a strong enterprise buy you know enterprises are it might be because it's good yeah I mean yeah it depends you know I don't know why Google has stumbled so much interesting yeah before we wrap up the AI related news I want to share a summary that risky business

[02:03:12] news produced from a much longer report by Bit Defender since I always try to go to the source I track down Bit Defender right up but it contains so much superfluous information that I returned just for risky businesses much shorter summary which still offers the information that we care about so here's what they

[02:03:45] to beef up its malware arsenal if this is a sign of things to come clustering threat actor behavior together for attribution purposes is about to get a lot more difficult by that they mean that you know all the security firms are able to do is guess at who did what based on the behavior that they see so they'll notice oh you know a lot

[02:04:15] of this code sure does look a lot like a lot of that code so these are probably related to each other somehow I mean you're literally left to kind of piece the forensics together that way and so what risky business starts by noting is that what this Chinese cyber espionage outfit is using its AI for is going to make this more

[02:04:44] difficult they said the Bit Defender report released last week describes seven remote access tool rat families all seven were created by a single cyber espionage actor Bitfender called Silk Parasite and five were previously undocumented meaning never been seen before the report authors have

[02:05:14] medium confidence that Silk Parasite is a China nexus actor targeting governments across Central Asia including Uzbekistan Turkmenistan and Kazakhstan back in November they say we wrote about what looked like an experiment to see how AI assisted hacking could support China's ministry of state security the approach those threat actors took at the time was to build an

[02:05:44] attack framework and let Claude do the hacking it was error prone and noisy and sometimes successful Silk Parasite by contrast is not using AI for YOLO hacking it is using it to support a cyber espionage to bit defender

[02:06:14] Silk Paradise develops tests debugs and iterates its own tooling maintains a structured build and deployment workflow and regularly rotates infrastructure encryption material payload names and persistence artifacts between deployments unquote they said rotating its infrastructure and indicators of compromise between deployments makes it harder to detect and link

[02:06:43] its activities together essentially these guys are using AI to become far more stealthy so that their individual instances of attack look like separate non commonly attributable attackers they wrote Silk Paradise also uses a variety of programming languages and command its seven

[02:07:13] different rats are written in different languages including .NET C++ Go or JavaScript command and control is accomplished by abusing Google Drive and internet communication protocols including HTML HTTP TCP DNS and TCP its malware also typically uses a modular plug-in architecture where additional functionality is only deployed when it's needed

[02:07:42] this means initial implants are relatively small and plug-ins are used to provide capabilities for clipboard monitoring key logging file management or interactive shell access this limits the exposure of the entire tool set during any single deployment it also allows Silk Parasite to update or replace individual components without having to change the entire implant and it minimizes the amount it writes to disk

[02:08:12] typically only the files required to get its malware up and running for victim organizations these measures make forensic analysis more complicated complete remediation is also more difficult once a particular implant is detected okay and remember that one of the key features that we're seeing now in cyber defense is the so-called IOCs the indicators of compromise but if every use

[02:08:41] has different indicators then nobody who's keeping track of all the latest seen you know previously seen indicators of compromise will have their alarms tripped when what looks like a brand new one-off attacker shows up so these guys are very cleverly using AI to basically maintain

[02:09:11] an amorphous presence and to be a chameleon attacker they said to us all the behaviors above are the hallmarks of a professional cyber espionage outfit of course doing all of this in a disciplined way is a lot of work and Bit Defender has evidence Silk Parasite is using AI to help it deliver

[02:09:41] this complex engineering yeah I mean this is a different scale of engineering that we've seen from bad guys before Silk Parasite's malware contains indicators of AI assisted development such as leftover test functions and placeholder encryption keys intriguingly implants the Bit Defender dubbed Goggin Rat and Nomad Rat share a high level

[02:10:10] architecture even though they are written in Go and C++ respectively and use different command and control protocols and code structure therefore BitDefender suspects that the same high level specification document meaning prompt was independently implemented twice with AI assistance BitDefender concedes that this structural similarity is not

[02:10:40] conclusive evidence but notes that it is the kind of thing an AI assisted workflow makes very easy this is the first example we've seen where the evidence tells a compelling story of a competent cyber espionage actor incorporating AI into its work practices Silk Parasite is taking the same disciplined approach to malware development and doing more of it it's creating more malware

[02:11:10] families to build redundancy making attribution and discovery harder and reduce the risk of compromise from any single exposure bitdefender has done a good job describing Silk Parasite's malware families and has published indicators of compromise that kind of exposure would once have set the group back significantly the group meaning Silk Parasite now that they figured out how to speed up their deployment work they'll be

[02:11:40] back better than ever relatively quickly then the discovery attribution and publication merry-go-round can start all over again so for me that's you know what's been reported here represents a sane entirely defensible and believable example of the way AI will be used to conduct offensive cyber operations apparently it's already happening certainly it will be in the future so

[02:12:09] much of what we're hearing about the coming AI driven you know cyber apocalypse when rogue AI agents freely roam the internet wreaking havoc wherever they choose none of that makes any sense to me what I expect to see is pretty much what we've been seeing only more broadly and deeper governments will use AI to infiltrate their espionage targets and criminal gangs will use it

[02:12:39] to infiltrate commercial enterprises then exfiltrate their valuable proprietary data and attempt to extort and blackmail them using the data they obtained I just don't see any coming cyber apocalypse you know yes everyone needs to be more wary and the probable targets you know of these operations governments and commercial enterprises storing data they must protect need to shore up their

[02:13:09] cyber defenses that's absolutely true you know pay more than passing lip service to their CISOs you know make everything as secure as you possibly can because we are going to see a much increased obvious increase in the use of AI okay this is the last of the two things I want to talk about

[02:13:39] and I'm excited about this because I think this is really interesting this is I think an interesting thought piece posted by one of as I said our favorite cryptographers Johns Hopkins professor Matthew Green a couple of weeks ago Matthew proposed something interesting that I think is worth pondering his blog's title was everything is about to go

[02:14:09] dark his piece explores the possible consequences of AI being used as it certainly is and is going to continue to be to make many systems far more secure nobody could argue that if Microsoft has just patched nearly a thousand bugs that Windows is

[02:14:39] Matthew suggests there may be unforeseen consequences of that so here's what he wrote he said I'm coming down from spending a few days at Usenik's security right here in my hometown of Baltimore this means that my days have been taken up with two kinds of conversation first explaining to my colleagues why Baltimore is not actually like the wire which of course is HBO's famous series

[02:15:10] he said and second trying not to talk about AI he said I'm going to break that second rule now he said I have this is Matthew Green saying I have many worries about what AI means for our field for various definitions of field but in this post I want to focus on just one thing I've started worrying about and it's a perverse thing

[02:15:40] specifically I'm concerned that AI is going to make software much too secure well that doesn't sound so bad on the surface there's a consequence to this he says he said I mean something very specific I'm concerned that US intelligence and law enforcement agencies are about to go dark meaning that they're

[02:16:10] going to suddenly lose a huge portion of their capability and that this isn't going to be a problem for those agencies but also for those of us who value computer security and privacy in general okay so what does going dark mean in the era of law enforcement hacking he said to explain how we got here we

[02:16:40] need to talk about recent history this actually gives me a real excuse to reference the wire just because it's a perfect snapshot of what electronic surveillance looked like way back in 2002 if if you've seen the first season you'll recall that it's about cops wiretapping drug dealers who use pay phones and burners burner phones yeah yep the mobile phones in the show are relatively new

[02:17:09] technology for the time but from a technological perspective nothing in this scenario would have shocked a cop who jumped forward from say 1989 the change began in the late 2000s thanks to the rise of smartphones and texting because smartphones can actually store data as well as conveying it the contents of those phones quickly became a useful new source

[02:17:39] of law enforcement capability or they were until 2010 when Apple began encrypting iPhone storage using a key derived from the user's passcode he says Android phones followed shortly thereafter the next year Apple deployed end-to-end encryption in iPhone text messages by 2014 a tiny texting startup named WhatsApp had gathered 600

[02:18:09] million users worldwide by 2016 those users now nearly a billion strong were all using default end-to-end encrypted messaging and calls these two trends the move from calls to texts and texts to encrypted data happened very rapidly the FBI and law enforcement agencies were not insensitive to what was happening

[02:18:38] in 2014 director Comey announced an national conversation about what providers could do or be compelled to do to make these new communications media legible to law enforcement and counter intelligence in 2016 the agency quit talking and took their theory

[02:19:08] to court when a terrorist attack left the FBI holding a shooter's locked iPhone the agency ordered Apple to give them access the company refused what broke the stalemate and to some extent ended the going dark conversation was something neither the FBI nor Apple expected an outside company announced that there was no need for Apple's assistance

[02:19:37] they could simply hack the phone the Apple versus FBI case has turned out to be a microcosm of the whole going dark debate for the next decade law enforcement and intelligence agencies continued to ask for exceptional access back doors but the urgency was gone agencies and manufacturers both knew that law enforcement could and would purchase targeted

[02:20:07] hacking tools like gray key for phone unlocking or even remote exploitation tools like NSO groups Pegasus if they needed them badly enough vendors like Apple and Google continued to play a vigorous defense closing vulnerabilities as soon as they learned about them but offensive vulnerability hunters consistently managed to keep the edge but

[02:20:37] now today there's a very good chance that all this is about to be history because the era of AI bug hunting is here this April just four months ago Anthropic announced a new model called Mythos that happened to be unusually skilled at software vulnerability discovery the US government temporarily blocked its export restricting access to

[02:21:07] US agencies and trusted vendors while the ban was dramatic and made for good PR I'm sorry while the ban was dramatic and made for good PR it turned out to be mostly pointless open AI along with Chinese open weight model labs like ZAI and Moonshot have since demonstrated that vulnerability finding is not something that a

[02:21:36] single lab is likely to hold a monopoly on the growing list of serious vulnerabilities these models have found is getting scarier and more impressive by the day at first glance this might seem like good news for the offensive team and for hackers in general he says but I doubt that's how this will play out in the long term defenders are now in the process of patching every bug they can find often

[02:22:06] decades worth of bugs and the backlog feels huge but they're making progress entire CI tool chains are being rebuilt to incorporate AI based vulnerability scanning before software ever reaches the point where a human will touch it while I doubt this means that every bug will be found in the real world it does mean it does feel likely

[02:22:36] that we're going to hit some sort of a ceiling on the number of useful bugs and we'll probably hit it soon so in this regard Matthew and I are in complete agreement we're going to see this bug discovery and patching rate eventually drop and drop near to zero he says thus over the next two years major pieces of software are

[02:23:06] likely to run out of remotely exploitable bugs he says obviously I think this is great but for law enforcement and offensive intelligence agencies it's going to be a nightmare for the first time since 2010 law enforcement might experience what it looks like to really go dark across a huge category of advanced well-maintained devices

[02:23:35] and pieces of software so how is this a problem the debate over exceptional access mechanisms never really went away in some places like the UK it even metastasized into something worse here in the US it mostly went into hibernation some of the slow down can legitimately be attributed to expert pushback academics and industry engineers

[02:24:04] pointing out the risk that back doors might be abused by the very adversaries that agencies are supposed to be protecting us against but I fear he writes that this was less of a principled pause and more of a market that was just pricing supply the destruction of the low hanging vulnerability fruit will make law enforcement and intelligence agencies needs much more

[02:24:34] acute the demand for constructed intentional back doors will restart in earnest the result will be enormous pressure on industry to re-architect their systems to make their systems amenable to exceptional access in some cases governments will ask for these capabilities in the expectation that they'll be useful for spying on other

[02:25:04] governments us strategy that might have been undetectable in the pre ai era but that probably will be less productive now the results are unpredictable one result might be that non us governments entirely remove their dependence on us software the worst part about this dynamic is that these potential new back doors will probably only affect the countries

[02:25:32] that demand them, meaning that they will be primarily useful for allowing the U.S. to weaken its own systems. This will in turn allow foreign adversaries to find new ways to attack our communications. This deliberate self-sabotage will happen just at a moment when we're finally getting a handle on securing our own infrastructure. So what do we do about it? He says, I honestly

[02:26:00] have no idea. This is not a call to action for experts to rally behind a sophisticated plan. Like so many things about the AI revolution, it's just occurring to me that we're on a long, greasy slide to a place that will look different than where we are today. Just realizing this doesn't mean I have any strategy in mind to avoid it. In this case, we're just going to have to hope

[02:26:30] that this time we make the right choices for no other reason than they're right. So I wanted to share this because I think it's a brilliant forward-looking take on our near-term future. Over here in the U.S., we've been watching the United Kingdom and the European Union wrestling

[02:26:55] wrestling with this issue over and over. And we've comfortably watched it refuse to die. Well, comfortable from a distance. Watched it refuse to die. They, you know, it just will not die. But this very issue may soon be visiting our shores here in the States. We've watched

[02:27:16] our United States enact laws that have forced adult content websites to blackout access across entire states because there's currently no practical way to guarantee the age of anyone visiting. And at this very moment, Utah's Senate Bill 73, which was signed into law on March 19th and which

[02:27:44] sailed through Utah's state government passing 22 to 2 in their Senate and 66 to 1 in their House, meaning a stunning majority, was written to take effect last Thursday, September 3rd, but was just extended by 14 days, two weeks to September 17th, Thursday after next. That onerous

[02:28:11] bill deems a person who is physically located in Utah to be a Utah user, regardless of the IP address they present to any age-restricted internet service. Its shorthand is the anti-VPN legislation,

[02:28:37] because it's clearly meant to curtail the use of VPNs and other proxies as a means of geo-relocating. No one has any idea what's going to happen there, but it's going to be interesting to watch, because, you know, basically they are trying to prevent adult age-restricted internet service websites from allowing connections from VPNs, but not all VPNs declare themselves as such.

[02:29:07] So my point is, there is a clear tension growing between the legislatively protected privacy rights of citizens, and in some cases we have constitutionally protected privacy rights, of citizens in many major democracies, and the perceived needs of their own governments

[02:29:33] to conditionally violate those rights. In an analog and pre-encrypted world, law enforcement and intelligence services were able to sneak around to get what they believed that they needed. And even after nominal encryption was emplaced, until the advent of AI, the large supply of latent bugs

[02:29:56] allowed these same agencies to ignore privacy whenever they felt it was necessary to do so. Matthew's point is, that fragile status quo is soon to end. What will replace it? Richard Danton Good question. He's a smart guy, and he's very AI-aware, too. He's actually

[02:30:25] a pretty pro-AI guy, so. Matthew Dixon Yeah, and he's also one of the guys that, you know, that the legislators pull into, you know, for Senate hearings to find out what he thinks. And, I mean, so, you know, here we're saying, look how much more secure Windows is today than it was yesterday with a thousand new problems fixed. Yeah. And, and we know that Android and iOS are

[02:30:54] going to be on and Mac OS are going to be on, they're all going to be on the same curve. They're all going to be getting the benefit of this. And we could argue that AI is going to help future errors, as Matt said, not get into production code. So we're going to fix this. And, you know, Apple has struggled to keep their phone from being hackable. It's been a struggle.

[02:31:19] They're probably going to win. And then what? Well, I mean, they've been complaining about going dark as you pointed out since James Comey. And yet digital technologies have given them more ways of seeing into our lives than ever before. Exactly. And, and, and people like the NSO group with Pegasus, Pegasus have been able to keep compromising people's phones. Yeah.

[02:31:49] What, what happens when that changes? When they can't. Well, then we're back to the way it used to be when, when law enforcement didn't know every darn thing that was going on inside your house, inside your mind. Except that we were, we were talking over analog phone lines and wondering if that little click and static sound with somebody listening. Like there's no question. Law enforcement is always going to want more, even if they weren't going dark, they're always going to be pushing for

[02:32:17] this, right? This is just, they have been, and they will continue to push for this. Well, the problem is our legislators are now going to, you know, they're going to come under pressure to, to, to, to legislate this. Right. I see this as part of a much larger uncertainty in general. We are entering into, it's really, I would say, fairly chaotic. We are in a time of, we are in a time of phenomenal change. Right. I agree. I mean,

[02:32:44] one thing we know about chaotic systems is very hard to predict. It's just not, they're not deterministic. And I think, I think this is chaos. We don't know what AI is going to do. We, it could, the scale ranges from, it's just more computer programming to it's a, it's an alien intelligence in our midst. And I don't know where it's going to land on that scale. Uh, and I don't know what the disruption is going to be. Will people lose their jobs? I'm not, I don't even know if that's

[02:33:12] clear. So it's all very, uh, through a glass darkly. So I don't, I just don't know if we can make any sensible plans, I guess is what I'm saying. We should be a, there's a storm a coming. Uh, and I don't know what you do to prepare for it, except maybe, uh, you know, get more rice. I don't know. Buckle up, buckle up. It's going to be a bumpy night. Okay. Last break. And then I

[02:33:40] am very excited to share this endless lifelong engineers observation about the danger of become overly reliant on automation. Yes. And we have a new, a new automation capability in town. AI. AI, AI, AI. That's what I have been saying. All right. We're going to get to that in just a bit

[02:34:06] before we do. Oops. I don't know why it's doing that. Let me turn that off. Thank you very much. Uh, restream is just giving me all sorts of fun before we do. I do want to talk a little bit about our club club twit, which makes this show and all the shows we do possible, uh, without club twit. Uh, we would have to cut back at least 30%, maybe 40%. That's a third of our shows, a third of our

[02:34:34] hosts, a third of our staff. Uh, and I don't want to do that. I like what we're doing. I hope you enjoy it. I think you're getting value out of it. Um, and if you are, I'd like to invite you to join the club. Here's the sad truth of the matter. Only about one in 15 people who listen to this show are club members. If we could get that to one in 10, we can get that to one. And actually it's not

[02:34:59] one in 15. What am I saying? It isn't even close to one in 15. It's about one in 50, maybe it's not even 2%. It's less than that. If we could get to one in 20, we would never have to worry about ads or anything else. We would be well supported. We could grow, we could have new shows. We could have a bigger variety of shows for you. Uh, I just, I feel like we're so close and your help really could

[02:35:28] put us over the top. I would like to not be beholden to advertising. I would like to be supported. We, I don't ever want to do a paywall. I think what we do is too important. I want to always give it away, but if the people who can, and I know you can't, you all can't afford 10 bucks a month. And if you can't, that's fine. But if the people can join the club, support what we're doing, we can go into the future. I think doing it even better. So that's my pitch.

[02:35:56] twit.tv slash club. Twit is the website. Everything we do is expensive. Running this website is expensive. Running these shows is expensive. Uh, paying our hosts and our, uh, our staff is expensive. Uh, it, it's not me. I often don't get paid at all. That's fine. Uh, but I do want everybody else to get paid. So please twit.tv slash club twit. You do get access to the club

[02:36:21] to a discord ad free versions of all the shows, chapter markers in all the shows. Uh, you get all that special programming. For instance, tomorrow we're going to cover the Apple keynote. We can't do that in public. Apple doesn't like that, but we can do it in the club. It's a private broadcast. So that's how we're going to do it. If you want to see those private shows, the only thing we do that's private twit.tv slash club twit. That's all. That's all I'm going to say. Let your conscience

[02:36:47] be your guide. No, I don't want you to feel guilty. If, if you can't afford it, if you can't do it, that's fine. But we would sure like to have you in the club. Uh, now on we go with Mr. G, a guest submitted article recently appeared in the IEEE spectrum publication. Uh, it captured my imagination and it feels very important to me. Uh, and I believe it's going to resonate deeply with

[02:37:12] many of this podcast listeners who've been around the block a few times. Um, for those who don't know, IEEE is the public is the abbreviation for the Institute of electrical and electronics engineers. Uh, it was founded as the AIEE, the American Institute of electrical engineers, believe it

[02:37:34] or not back in 1884, an astonishing 142 years ago. Um, and then it was later renamed to IEEE after it's 1912 merger with the Institute of radio engineers, the IRE. Um, and you know, to put the Institute's age into perspective, it was at the start of the year of its founding

[02:37:56] that a penniless genius by the name of Nikola Tesla arrived in New York to work for an already famous industrialist by the name of Thomas Alva Edison. But anyway, I digress. Uh, today the IEEE describes itself as the world's largest technical professional organization dedicated to

[02:38:20] advancing technology for the benefit of humanity. And the article I encountered, which so galvanized me carried the headline, AI efficiency could cost us the next generation of experts. And the articles, the articles teaser read lessons from aviation and nuclear power show how to preserve human skills.

[02:38:44] Uh, and as I said, I, I believe what's said here is extremely important. So, so see what you think. It's author wrote a little over a decade ago, I led the controls design for a first of its kind full digital control system for a U S nuclear plant. It was on paper, a beautiful machine engineered to

[02:39:12] run itself the way a modern airliner does with operators watching over a system that rarely needed them. And we made a decision that to an efficiency minded observer looked backward. We deliberately left manual steps inside sequences. The system could execute on its own. We were solving a specific problem.

[02:39:39] An operator who only ever supervises automation slowly stops being an operator. The hands go cold. The mental model of what the plant is actually doing gets fuzzy. Then comes the day the automation hands control back. It's always the worst day because automation only quits when it's confused or in

[02:40:07] trouble. But by then you have a person in the chair who has not truly operated the thing in years. The manual steps we included in its design were there to keep the human current. It was inefficient by design on purpose. The plant, as it happened, was never built. It was shelved amid the politics and economics that surround

[02:40:36] nuclear power in this country for reasons that had nothing to do with the engineering. But the design instinct outlived the project. And I've become, he wrote, to, I've come to believe it's the most useful idea I can offer to the argument now consuming every boardroom. What happens to human expertise when AI

[02:41:01] does the work that used to build it? The data has become hard to wave away. A Harvard University working paper covering some 65 million workers and more than 280,000 US firms found that after companies adopted generative AI, junior employment, junior employment fell roughly 9% within six quarters relative to non-adopters,

[02:41:31] while senior employment kept right on growing. A Stanford analysis of ADP payroll records points the same way. The youngest workers in the most AI-exposed occupations lost ground after late 2022, while their more experienced colleagues held theirs. The Stanford researchers found that the losses

[02:41:55] concentrate where AI automates the work, where it merely augments junior employment holds steady or rises. The causal story is still contested and honestly requires saying so. Researchers at the New York Fed attribute much of the rise in young graduate unemployment not to AI, but to remote work, arguing that firms are

[02:42:20] reluctant to hire inexperienced people whom they cannot train and mentor at a distance. But notice what the explanations share. Whether a model is absorbing the formative work or distance is severing the mentorship around it, both describe the same broken mechanism. The apprenticeship channel through which expertise

[02:42:43] passes from senior to junior. Either way, entry level has quietly come to mean three years of experience required. To strip away the noise, or he says, strip away the noise and you're left with one deceptively simple problem.

[02:43:02] You cannot become a senior engineer without first being a junior one. Expertise is not downloaded. It is earned through failed builds, dead-end debugging sessions, and the why on earth did that work moments that a capable AI will now happily spare the newcomer. Spare them enough of those,

[02:43:28] and you produce a cohort that can supervise a model on paper, but never developed the gut sense to know when the model is confidently, catastrophically wrong.

[02:43:41] Most of the commentary stops at the diagnosis or reaches for policy solutions that treat the loss of junior jobs as an economic problem. Yet it's also an engineering problem. And safety-critical fields have already spent decades learning how to solve it. He said,

[02:44:02] My own career began at the sharp end of automation. My first job out of school was verifying and validating the software in digital jet engine controller that decides faster than any pilot could how a fighter plane's engine responds.

[02:44:25] Even then, in the late 1980s, the central tension was visible. The machine outperforms the human in routine cases, but the human is all that stands between the aircraft and disaster in the cases the machine did not anticipate.

[02:44:43] This tension is known as the automation paradox, in which increasingly capable automation gives human operators less practice while leaving them with only the most difficult situations. Aviation learned repeatedly and expensively what happens when human skills atrophy inside that gap.

[02:45:12] The canonical example is Air France Flight 447, which fell into the Atlantic in 2009. The proximity of the plane was mundane. Iced over airspeed sensors fed the autopilot bad data, and it did what it is designed to do. It disconnected and handed control of the airplane back to the crew.

[02:45:40] What followed was not a hardware failure. It was a competence failure. A recoverable situation became an unrecoverable one because the pilots, conditioned by thousands of hours of watching the automation fly, could not read a high-altitude aerodynamic stall and hand fly their way out of it.

[02:46:06] The airplane was working. The airplane was working. The training the automation had quietly eroded was not. The industry's response is instructive, and it's the same move we made in that nuclear control room. It did not rip out the autopilot, but built deliberate manual practice back in.

[02:46:30] In 2017, the FAA issued Safety Alert for Operators 17007, Manual Flight Operations Proficiency, declaring that manual flight is the foundation upon which other technical flying skills are built. The alert formally recognized skill decay as a hazard in its own right.

[02:46:56] Some airlines amended their procedures to encourage hand flying both the initial climb and initial descent in benign conditions, knowingly trading a sliver of fuel efficiency to keep the crew's raw flying skills alive. That trade is the whole point. A perfectly optimized system that produces incompetent operators is not optimized at all.

[02:47:24] It has simply moved its failure mode somewhere the spreadsheet cannot see. Put the aviation lesson and the nuclear instinct side by side, and they point to one design pattern we now need in AI augmented work. The deliberate manual gate.

[02:47:44] A manual gate is a point in a workflow where a human takes the controls, not because it is the fastest way to get the task done, and not as a safety interlock, but specifically to exercise and preserve a skill that would otherwise decay. The distinguishing feature is that it is chosen.

[02:48:08] You decide as a matter of design which competencies your organization must keep alive in human beings because those are the ones you will need on the bad day. When you engineer the friction required to keep them warm. Picture how this might work on a software team that leans on AI for most of its code.

[02:48:35] The team places a manual gate around the skill it can least afford to lose. Debugging. When a defect surfaces in a critical module, the assigned engineer, deliberately, often a junior one, must first reproduce the failure, trace it to root cause, and write an automated test that captures the bug.

[02:49:00] All without the AI assistant switched off. Only after the engineer commits to a diagnosis does the model come back on to propose the fix, generate alternatives, and sweep the code base for similar bugs. The engineer then compares their diagnosis against the models.

[02:49:28] When the two disagree, that's the design working, surfacing the disagreement before the bad day instead of during it. It's also the design teaching and training the junior engineer. This approach reframes the junior engineer entirely. The instinct today is to let AI do the entry-level work because it is faster, cheaper, and capable.

[02:49:57] But some of that work is not overhead to be eliminated. It is the training apparatus of your future senior staff. And you should protect it the way you'd protect any other piece of critical infrastructure. It may not be efficient this quarter, but dismantling it quietly mortgages your capability a decade away.

[02:50:21] None of this is free, and pretending otherwise would insult the people who have to sign the budgets. A deliberate manual gate is, by construction, less efficient in the near term than full automation. Keeping juniors doing formative work and running the manual sequences costs something now to protect something later.

[02:50:46] That's a hard sell in a market that judges most leaders on quarterly results. A hired executive who carries, he has an air quotes, unnecessary humans that AI could replace, will hear about it from the board long before the payoff arrives. The math only works for someone insulated from that pressure. A founder with control. A private company.

[02:51:13] An institution with a genuinely long horizon. Or a regulator willing to require workers to demonstrate their skills regularly as pilots must. This all means the organizations most likely to preserve their own expertise are the ones structurally able to spend short-term margin on long-term capability. Everyone else will need a push from the outside.

[02:51:42] So here's the argument in one line. Deliberate inefficiency is not waste. In safety critical engineering, we have always known it as insurance. And we buy it on purpose. As AI takes over the work where expertise is forged, the smart move is not to resist the automation.

[02:52:07] It is to keep our hands on the controls by design so that when the automation fails, as it always eventually does, there's still someone in the chair who knows how to fly. So I think that's a fantastic piece of well-reasoned engineering.

[02:52:29] And I would recommend it without hesitation to anyone's boss who may currently be enraptured by AI's truly deliverable ability to eliminate all of those pesky junior engineering jobs. You know, we already had stunningly good AI with GPT-5.6 Fable and Mythos.

[02:52:51] And though it's going to take some time for the world to fully assess what OpenAI has just given us with GPT-6 Astra, the early take is that it currently, you know, outperforms some of the other frontier models, you know, from OpenAI, Anthropic, and Google. Whether or not and to what degree that's true, you know, and even the fact that this is nothing more than a snapshot in time.

[02:53:19] That's important because, of course, this wasn't true a week ago, but maybe true today. And we know that we are just at the beginning of all this. We can feel how rapidly this field is changing. And so that's my point. If we've learned anything, it's that no one will ever have a long-term defensible position in AI. Sure, you know, they're going to be bouncing back and forth.

[02:53:48] And Leo, you already knew that Anthropic was, what, a week away or so from, you know, releasing their next model in this, you know, never ending series. Supposedly. I mean, it's rumored anyway that they're ready with Fable 5. So, yeah. Anyway, I think without the deliberate creation of an environment that is designed to nurture and mature junior developers,

[02:54:15] organizations will be left completely dependent upon automation for mission-critical functions. I mean, I know that coders who come out of university with no experience, they're just not useful. They actually can't do anything. And you could argue that they're going to be a lot less useful. All they'll know how to do is drive AI. No coding. I don't even know if programming will be taught in the future. It'll be prompt engineering. Right.

[02:54:48] And you think that we need to know how to code? We probably need to know how to read what AI does when it writes a bug. You know, can we, I don't know yet whether we're going to be able to say to AI, fix the bug which you created. Right. Maybe we will.

[02:55:16] I mean, so I guess the question is, okay, so certainly there will, well, AI is also coding itself now, isn't it? Yeah. So maybe it will be a completely lost art except here in Irvine. Yeah. Yeah, right. There'll be a few people. I mean, I think there'll always be a few people who code by hand, just like there are a few people who make their own furniture. But, um. And play the piano. Yeah.

[02:55:46] Instead of, you know, turning on Apple Music. Yeah. I just, I don't know, I don't, code is a funny thing because code is getting the computer to do something by translating your English language thoughts into machine code. Of course, when you're kids, you're writing most of the machine code. Your intent. Even unvoiced, your intent. Because it might be unvoiced. Right.

[02:56:10] I mean, all, and as we've moved higher up, you still, you do low-level coding, but most people are doing high-level coding, which isn't really that closely related to what's going on in the computer. I guess there are, there is a case to be made for people who need to be able to look at the traces and see what's actually happening. Well, for a long time, compilers had bugs.

[02:56:33] And which meant you would, you would properly express yourself in the high-level language and you still wouldn't get the program doing what you told it to do. Right. Right. Which, to me, is another case of, we don't know what the hell's going to happen. Well, there are people talking about AI slop code and that we're now talking, that we're going to be introducing a whole new generation of bugs. I don't know one way or the other. I think that's old farts talking.

[02:57:04] Could well be. Yeah. I think that that's people who are very tied to the old way of doing things. Slop is a pejorative that I don't think is completely fair. It doesn't code the same way you do. That's true. But the thing about computer programs is, you know, if they're correct, they work. Problem is these little edge cases where it works most of the time. It doesn't work all the time.

[02:57:30] But I think a computer is ultimately, at least I don't see any reason why a computer wouldn't ultimately be the best at figuring out what's going on. It's, right? It's better at it than we are. I was arguably one of the early people to say that AI was going to be extremely good at code. This is what it does. Exactly. It understands. So really what the AI is doing that's hard for the AI is, is it good at us?

[02:58:00] And I think most of the time when the AI fails us is because it's not as good at us or we're not as good at it. It's a communication problem between us and the AI, a misunderstanding or miscommunication or misdirection. I mean, I literally haven't looked at code in a long time.

[02:58:19] And I think most many of the people like Darren, who's a very proficient coder in our discord, I looked up to him because, you know, he's the guy who would finish all the advent of code problems lickety split. So he's a very, very accomplished programmer, done it for years, his whole life. He says, I haven't coded anything myself for like a year. Wow.

[02:58:41] I don't think he, and I think many of the most proficient coders I know, people like David Hennemeyer Hansen, don't look at the code. Now, his operating system on Markey, which is a version of Linux that's entirely AI coded, is kind of wonky. But that doesn't, I'm not sure I blame the AI.

[02:59:03] I mean, many of the coders who are using AI at this point have stopped looking at the code. So I don't know what that means. I don't know. I think, and Darren's saying, you can absolutely build great quality software without knowing anything about writing code right now. And I think that's only going to get better. But it's absolutely true.

[02:59:32] There needs to be an entry level positions. You know, I think the analogy is to the business I'm in of broadcasting. You're terrible when you start. You really are. You're not good at it. And you need somewhere to start. You need, basically, it took me a year to find a radio station that was so tacky and small that it would let me learn how to be a broadcaster there.

[02:59:59] And they put me on the middle of the night, midnight to 6 a.m. on Sunday morning, because that was where I could do the least harm. But I got a chance to learn. And it's, you know, we know it's called apprenticeship. Right. The idea is you apprentice with a master and you learn his art. Right. But then I look at my son, who really has no broadcast training, who just did it all in public on TikTok and has done a little bit, done pretty well for himself.

[03:00:29] Well, better than his dad ever did. And he did it, by the way, in a hundredth of the time his dad did. I mean, this is literally, he's gone from nothing to total success in whatever this business is, sort of like broadcasting, in a couple of years. Took me 20 years to, 10 years to get halfway decent. So it's a different time.

[03:00:53] And I think a lot of what you're seeing in a lot of this complaining and conveiling about AI is us, the older, our older generation going, well, that's not how you do it. My day. You had to hold the hammer in your hand. How do you know what's in the registers? Yeah, you got to know what's in your registers. Do you care what's in your registries? You don't, you do. And you know intimately. But I don't know if you need to. And certainly somebody who's programming in C++ probably doesn't.

[03:01:23] No, you don't have any access unless you explicitly tell it you want. Yeah, you don't really care. Get a bit of MASM. You don't care. So I don't, I just don't know. No, it's, I am. We're an unknown, uncharted territory. We are riding a beast. We are, it's a tsunami headed for us. Yeah. And there are a number of people who are going run. And the number of people are going, oh, look, the water is going out. I see seashells. And I think I might be in that latter group.

[03:01:52] And the wave is going to hit me. Steve, always, always provocative, intriguing, entertaining, and informative. It's a great show. And I really appreciate you putting all this work in. You do every week. In fact, you can subscribe to his show notes, 21 pages this week of really good stuff. Read along with the show. Read it yourself. There's links. There's images. It's everything you need. It's basically a book every week. Have you ever counted the number of words?

[03:02:22] Is it three or four thousand? Read it by going to grc.com slash email. There's two things you'll accomplish by going there. You can give him your email and he will whitelist it so you can email him, which is really nice. Send him pictures of the week, ideas, questions, comments, suggestions. But below that, there's two checkboxes.

[03:02:47] One for the newsletter, the weekly mailing of the show notes. Another for a very infrequent mailing of new products. Speaking of products, Steve has, of course, Spinrite, the world's best mass storage maintenance, recovery, and enhancing utility. And you can get it right now at grc.com. That's Steve's bread and butter. Handwritten. He knows every single register in that program. He knows what it's doing at all times, right? There's no mystery.

[03:03:17] You know exactly what's happening in that machine, which is kind of amazing. And it's really good as a result. He also does a lovely little $10 program called the DNS Benchmark Pro. So you can get the fastest DNS. You know, AI can't do that for you. You have to have a deterministic program like Benchmark. You have to. Somebody has to write a program to do that. And Steve did. Thanks to Steve. He has the show as well.

[03:03:45] He has 16 kilobit audio and 64 kilobit audio. He also has wonderful transcriptions written by a human being, Elaine Ferris. They come in usually a couple of days after the show. All of that at grc.com. We have audio and video. Somewhat larger audio file. And video as well. You can download that from twit.tv slash sn. There is a YouTube channel dedicated to it. And actually the best thing to do is subscribe and your favorite podcast client. So you just get it automatically.

[03:04:14] You don't have to think about it. We record live every Tuesday right after Mac break weekly. That's 1.30 Pacific, 4.30 Eastern, 20.30 UTC. And you can watch us do it live. If you want the absolute unedited best version of the show, all the expletives, all the nudity. It's all there. Computer not working and rebooting. Yeah, mostly that. If you watch live, there's several places you can do that. Of course, if you're in the club, and I hope you are, the Discord.

[03:04:45] But you can also watch on, everybody can. YouTube, Twitch, X, Facebook, LinkedIn, or Kik. And if you're chatting there, I can see the chat. So we can chat back and forth. I often do that. Steve, that's it for us. I will see you next week. And maybe I'll have a new iPhone. No, I won't have it by then. Till then. And we'll have a breakdown of Patch Tuesdays. Fantastic patch count. And we'll know what Apple has wrought.

[03:05:14] Holy camoly. Thanks, Steve. Thanks, everybody. We'll see you next time on Security Now. Hey, everybody. The show's over, but the listening continues at twit.tv. In fact, we've got so many great shows, I don't even know where to start. How about Intelligent Machines? That's the show where we cover AI, robotics, and all the smart stuff all around you. With two of the most fun, smart people I know, Paris Martineau and Jeff Jarvis. It's often philosophic. We sometimes debate.

[03:05:42] But if you're interested in AI in this world that is changing so fast, this is the show for you. Intelligent Machines. You'll find it at twit.tv slash I am our website or wherever you get your podcasts. I'll be looking for you.

Firefox smart window,AI-driven expertise loss,Microsoft Patch Tuesday, Windows security updates, OpenAI GPT-6 Astra, NVIDIA Hugging Face acquisition, hypervisor-protected code integrity, memory integrity, CISA cybersecurity services,