All right, thank you so much. It's really nice to be with everybody. To set a little expectations, my goal is to keep this into the 30, 35 minute range. So I'll try to keep us on pace here, but I'm really excited to hear from you, BJ, and kind of take us through your experience with deploying AI into a SOC in the real world, kind of the pitfalls and the issues that you've run into. I think our audience will really benefit from the perspective of somebody who's actually done it. I know we're all talking about it. We all sort of are hearing it in the news and getting beaten with it at the conferences, but it's rare to actually get to interact with somebody who's walked the walk. So, AI anxiety is real, 100%. We all understand, you know, the threat is agentic, but take us through, BJ, how you get over the AI anxiety dynamic, kind of how you see it playing out in the SOC.
That's an awesome starting point. The anxiety is real. It's palpable. You know, I think one of the key things that really set the right pace and really helped me understand the anxiety we're dealing with is, I was presenting to a board one time and they had a great question, and I've had it more than once, which is: how do you know you're spending the time on what you should be spending the time on? And that's the situation. We've got a massive amount of alerts. This alert fatigue has been real for many years and we've been fighting it for the longest time, and it's such a basic question of how do you know you're spending time on what you should be spending time on.
That's where I think the agentic SOC is really going to help us get towards some better fidelity and accuracy and really build that confidence that we know we're spending time on the right things, because we look at the amount of data that we have that just gets unused or not looked at. One of the other things I've always had as a risk is, what if we miss something? What if we don't look at the right thing? But as we really start to bring AI and the machine capabilities to the table, we get much better visibility. We get more coverage, more execution, and, as we're going to talk about a little bit today, we get to uplevel our resources as well.
So, the anxiety's there, it's palpable, but we can work through it. A couple things as we do so: you've got to understand that things can go wrong. You've got to know what those are. You've got to work through them, but you also have to start to think about what is your success and how do you get there. Because at the end of the day, we're not changing the cyber landscape too significantly. We are changing some facets of how we do it. But bad guys are here. They're attacking the good guys. We're here to defend. So those logistics still remain.
How do you deal with it in your mind? What I run into a lot is people want to talk about hallucinations and how do you know the agents are doing the right thing. There's just a whole level of trust that people seem to need to get over, or at least acclimate themselves to, before they're even willing to go down this road. How did you deal with that conceptually?
Somebody had made a good comment to me one time. They said, "Look, your agentic SOC is like bringing a load of interns into your company. You need that starting point." When you look at it like that, your agentic SOC is going to come in and it's going to need some guidance. It's going to need some support. You need to point it in the right direction, help it, push it towards what you want it to do. And then it can execute and do some amazing things. I've had some of the coolest deliverables come out of an intern. You give them the opportunity to run, and agentic is the same way. That anxiety is there, but understand what it is and how it works, coach it, guide it, get in and tweak and tune it, and you'll see some amazing stuff come out of it. Don't be afraid of it. You just can't be, because the reality is we have to leverage the agentic SOC because it is a machine versus machine fight out there. Humans versus humans, that's over. It's machine versus machine and we've got to be ready.
Yeah. It's almost like if you don't do it, you're fighting an asymmetric fight that you simply can't win.
You're just not going to. Yeah.
Yeah. So, listen, everybody out there who's been to a conference in the last six months knows. I walked the floor at Black Hat and you knew you were in the agentic SOC hype cycle. It's just crazy. Everybody and their brother is out there peddling this stuff. How do you wade through it? How do you keep in check what these guys are selling, because they can get very aggressive with the claims and the whole "the world is going to be solved by AI" kind of marketing.
Oh, the hype. It's so true. I was at Black Hat as well and you couldn't turn in any direction without seeing an agentic SOC this and an agentic SOC that. Take five steps and you run into five of them. It was just insane. But you do need to weed through the marketing hype. There's a lot out there. For me, I look for some red flags. When it seems too easy, it probably is. If it's too good to be true, it is too good to be true. SOC in a box where it's all black box? No, that's not going to do it. Fully automated, you don't have to touch anything or do anything, it just comes out and works? No, I don't believe that. That's snake oil. There are some snake oil agentic SOCs out there and you've got to be careful and watch out for them.
I also have a big red flag when they come in and say, "Hey, we're going to replace your entire team." No, I don't believe that. You're going to hear me say it over and over: an agentic SOC is not a replacement of people. It's an augmentation. It's an enrichment. It lets you uplevel your people. And the other thing to be careful of: you get all these ones that are promising everything, like, "Hey, we're going to solve everything for you straight out of the box," but you also have to look at the other ones that say, "Okay, we're going to do a lot, but you're going to have to add a lot of guardrails." You almost have to balance it. Are they fully into that? Are they more of a startup? Are they built in? So you've got to balance both sides of where the maturity of that vendor is. But don't believe the easiest one. Don't believe it's going to solve the world for you straight up. It's just not going to do it.
I'll say, just sort of riffing off of that, I don't think there are any headcount savings to be had in a well-run SOC team. You can upskill them, but the adversary has taken such an aggressive step forward that the idea that you're going to just rip and replace your existing SOC team with a bunch of AI agents is nuts. There's just too much work there. Maybe in the NOC, maybe in ITSM. I'm not seeing it in the SOC right now.
100%, I agree with you. And I'd go so far as to put a cautionary tale out to some of my peers in the industry, because the hard part is you're always fighting for budget. I've talked to a couple people who take the gamble and say, "Hey, we're going to put in an agentic SOC. If we get the budget, we can reduce headcount." That's not a wise gamble. I didn't do it that way. There's a value add. There's an augmentation and enrichment of your staff. But if your gamble is a headcount swap, you're going to be unpleasantly surprised.
Yeah, I do sort of repeat the joke, at least in terms of the hype cycle at Black Hat, that I left my booth and spent the rest of the conference accidentally at other people's booths. It's so hard to tell the difference between everybody out there.
So true.
So, listen, I think this is really relevant because it comes back to the SOC in a box nonsense, but what even is an agentic SOC? What are you actually looking for it to accomplish? We're both on the same page. We're not looking to automate our SOC team out. Again, maybe NOC, maybe ITSM, but definitely not SOC, at least not in the near term. So what are you trying to get out of it, in your experience? What's the value you were trying to harvest out of it?
I'm going to go back to that key question I started the conversation with, from the board: how do you know you're spending your time on the right thing? For me, it's an augmentation and uplift of the SOC staff. So we look at a couple different things. Alert fatigue is real. Alerts are all over the place. And that's another warning about some of the agentic SOCs: if all they are is another blinky light, you don't need more blinky lights. What you do need is greater, comprehensive coverage, being able to see all of that data.
No human's going to be able to consume all that data at once. For years we've been playing this fight of, I've got too much storage, the price is high, I'm going to have to cut this back. I've got too many alerts, I can't see it all. How do I bubble up? How do I prioritize? But we're at this point where the data is there, and you toss an AI at it and it's going to see it all, process it all, consume it all. In fact, it's smarter when you throw a lot of data at it. So for me it was, okay, make sure we can bring in the data. If I'm talking to a vendor and they're saying, "Hey, we're going to have to slowly feed your agentic SOC," that's the wrong conversation. You're going to need to feed it and you're going to need to do some training, tweaking, tuning, but let it consume the data. So you get better visibility, better coverage. You've got so much data, especially in larger enterprises. I've been at hundreds of thousands of nodes on the network. No human's going to analyze that. Apply your AI to let it do it. So first off, greater, wider comprehension and coverage. That was a big win.
The other thing is you've got to uplevel your staff. We always talk a lot in security operations about tier one, tier two, tier three, tier four. Tier one we look at as the entry level. That's your interns, your entry-level jobs, and they're doing the low-hanging fruit. I've seen it in two different conversations: there's a tier zero, but really tier zero is your tier one being upleveled to tier two. When you look at the amount of work they're doing, let your agentic SOC come in and play that tier one. Whether you call it tier zero or tier one is up to your own nomenclature. But the reality is your tier one becomes a level two. They can focus on more thought-provoking investigations. And it helps with burnout too, because people enjoy digging in and doing better investigations. Rinse and repeat over time, you get burnout. Let agentic do it.
That's where we get to that key point: you've got to let agentic act. Let it bring the data together. Let it put some information together. Let it help augment and level up all your staff. Now, the endgame, and we'll talk more about the goal, is to let it do more and more automation. But while you're building that trust, keep the human in the loop. When you uplevel your tier one to tier two, you're giving them the data to make better cognitive reasoning and ask more thought-provoking questions. That way, when they're doing the human in the loop, they're focused more on where they need to be. As you keep that rinse and repeat over time, you get further into agentic automation. So there's a balance there, and we'll talk more about that. But those are the key things I was looking for. Got to get to action.
The last thing I'll say on this is, as I met with my senior leadership running the SOC, I always had a balance, which is: don't tell me what you're going to do, tell me what you've been doing. It's this transition of thought to say we've got to let the execution occur, and enable the leadership, and eventually even enable the agentic to go execute on its own. It takes a little bit of time to get there, but if you keep that mindset, it's a good target to aim for.
One thing I've seen that's really encouraging about these systems, at least the well-developed agentic systems that combine all three of these elements, is their ability to upskill, to actually teach that L0 and get them to a productive L2 really quickly. It's really impressive, because if you watch the agent do the work and it's aggregating all this information, and it's not a black box system, it's actually an impressive teaching tool. So you can get that migration much faster than you otherwise would.
Listen, we all love this quote: "Be curious, not judgmental." Particularly when it was delivered in the darts scene. Fantastic. But tell me about the mindset going into this. What leads to the best result?
There are two main facets of it. One, you've got to be open-minded. Like you said, I love this quote, "Be curious, not judgmental." I'm a big Ted Lasso fan. If you haven't watched the show, you probably should. You've got to be open-minded. You've got to ask questions. Be inquisitive. Be curious. If you're just going to turn it on and let it go, you're going to realize it's not doing what you want it to. You're not going to get the value out. You're going to be judging, going, "Ah, the vendor didn't build me a good product." That's not the right way to go. What you need to do is come in and start asking the questions. What can it do? What can it do now? What can it do in the future? And how is it doing it?
I think one of the big questions I learned to ask was, how did it come to that conclusion? I could use four different AI models and ask the same question, and I'll get four different directional avenues to that response. The answers are probably very similar, but how it got there is going to be totally different. One of the things I learned when we were doing this: we had a couple scenarios we were running and we asked, "Well, how did you get to that reasoning?" and started to realize we had data signals we weren't leveraging in the past.
Humans become very dependent on certain data signals. If you've ever run a SOC, the number one, I'm going to call it almost a crutch, is your EDR. It's very valuable and gives you a lot of data, but then your SOC becomes so dependent on EDR, EDR, EDR. When we started to run a couple scenarios and asked, "How did you get to this?" we found some situations where EDR wasn't even in the loop. We started to realize things like, your identity, wow, that's a good source of data. You start to bring in your proxy logs. Those are good sources of data. Your data transitions. You start to bring in others and you get wider visibility. Going back to that risk I talked about earlier, making sure you don't miss anything: if you're solely dependent on your EDR, you're really blinding yourself and constraining your vision.
As we started to ask, "How did you get there? Why did you do that? What was the reasoning?" it opened our eyes to all these different signals. Then we asked the same question about, "Well, how would you solve it?" Now we started to see protective controls that were even better. Some were education-based, working with HR on new training. Some were firmer controls, either manual or automated. You start to look at things beyond EDR and firewall: your identity, your accounts, your ephemeral accounts, your non-human accounts, all these different ways to solve the problem. We didn't always see it the same way. So you've got to bring that curiosity factor and you'll be amazed at what avenues it opens up.
It's funny, because BJ, I know you go back into the forensics industry, as do I. I remember when we were pre-EDR and we were getting stuff off of network forensics. We were putting things together the best we could and you could get really far, and then EDR really changed the landscape. It's impressive stuff, it's good stuff, don't get me wrong, but it's taken people's eye off the ball in some respects, certainly for certain vectors of attack. And I think there's real value in what you're saying here, which is don't let it be too much of a crutch.
I agree. And I like that you brought up forensics. One thing I'll add on that: we used to always say there's no coincidence in forensics. The facts are there, but you had to be creative in how you brought the evidence together. There was one I used to talk about all the time, which was relativity by locality. If two pieces of evidence were close or nearby, then they're probably related. One test we were running as we went through some of the agentic SOC work: you always look at the absolute timestamp, that it has to happen at the exact same time. But we started to ask, well, what if they happen close to each other, off by 5 or 10 minutes? The AI looks at an absolute, but when we brought in the concept of giving yourself about a 10 to 15 minute buffer in the correlation of alerts, we started to see so much more. It reminded me of those forensics days. It doesn't have to be so perfect. Sometimes there's that nearness, and we started to get even higher KPIs and KRIs when we looked at that on fidelity.
Yeah, that makes total sense. It's harkening back to the low and slow hacker. Piecing together those investigations was always fascinating, and you definitely use the logic you were just presenting. So let's press ahead. This is a critical question, because there are as many naysayers as there are evangelists of this stuff, so getting this part right is critical. What does success look like? How do you recommend people break down that question?
First and foremost, it's going to be an attitude, which is: jump in, get your hands dirty, get your feet wet. Especially if you're a CISO or in leadership, understand what's going on. If you are simply trusting without understanding what it's doing, and assuming that it's always accurate, you're opening yourself up for errors and you're going to have an impact. So get in there. Let your teams get dirty. Work well with your partner, with the vendor. But take ownership of your environment. It's your environment. You've got to be there.
I like what it says here in the takeaway, which is success isn't accidental. You've got to get in there and be doing it. As you jump in feet first and start taking it on, setting up your use cases, understanding what it looks like, identifying hallucinations, which are still pretty low but do exist, you're really embracing the whole methodology and approach. That's another critical piece. You've got to get your hands dirty. You've got to test it. You've got to define it. You've got to come up with where you want to go. But also understand that agentic is going to make some errors. Like I said earlier, treat it like an intern. Treat it like your workforce. Question it, guide it. You've got procedures. You've got workflows. Leverage those, but treat it like that workforce. Stay curious. Stay on top of it. Watch it. Don't just plug it in and assume it's going to work. That's just asking for issues and impact, and you're increasing your risk.
But don't be afraid either. There's really no reason to be, because the phenomenal thing about these technologies is the barrier to entry, in terms of using the tool and getting value out of it, has dropped to just about zero. When we implement something, all the way down and up the HR stack, people can embrace it at whatever level they're comfortable with. It could be a CISO looking for dashboards and general questions about the environment. It could be an L1 practitioner trying to figure out what an alert is. So it really offers a spectrum of usage that previous tools didn't. It used to be you needed to know the layout of the UI, the workflow, the concepts ingrained in the software, and that's just not the case anymore.
Or you had to know the data structure, and that everything was parsed and organized. Now that's seamless. Now you're asking smarter questions: How do we get there? How do we stop it? Is it somewhere else in the environment? That's a big question. Is this somewhere else? And getting those answers fast. That's foundational.
And that really does lead to this slide, which is: how do we validate it? How do you get comfort? This is probably the number one question I get asked. After implementing, after it's up and running, after you're comfortable with it, how do you validate it? How do you make sure that over the last month it hasn't just ignored a bunch of alerts? How do you stay vigilant so you're not asleep at the wheel, and it's delivering continuous long-term results for you?
Oh, I love this. This is a great slide. I'm going to preface it a little bit, because we've all heard the term "trust but verify," and I changed this one specifically to say "trust but validate." Here's why. One of the main things you need to remember is you've got to get to a point of trust. A lot of people think I'm always talking about trusting the agent, but that's not necessarily the case. You've got to get the trust of the business. As you get to agentic, the business has to trust you to be able to let you use agentic. You have to trust that the agent is doing what you want it to do and what it should do, and that the output is correct. But as you do that, you're also building trust with the business, and the business says yes, we can let this automation succeed.
Because if the automation fails more than the adversary wins in an attack, basically, if you're causing more damage to your own company through your own inability to program your own agentic, you're a greater risk than the risk of a cyber attack, of a machine attacking you. And you don't want that. So you've got to know how to define your outcomes. Look at what you think it'll be. Watch the data. Watch the inputs. Watch the outputs. It will fail. Let's be clear. It's going to fail. But those fails should be considered wins, because you're going to see where the errors are. You're going to tune it. You're going to make it better. Hopefully your fails are in non-impactful or minimally impactful areas, but it's going to fail. Accept that, learn from it, and then go test it. Beat it up. Don't assume that all the answers are correct. Test. Validate against hallucination.
If it quarantines a system, what was the reason it decided to quarantine? If it quarantines a critical system, it's probably a fail. Could be. If you're causing a business impact, figure out why it came to that conclusion. Figure out if it has its own biases built in. This is your chance to really test it and build that trust. If you've ever heard conversations on trust, we all know it's very easy to lose trust and hard to build it. It takes one failure to easily lose that trust. It takes eight, nine, ten successes to even incrementally grow back that trust. So make sure you're methodical in how you do this.
Yeah, I think that's incredible advice. And you're talking to somebody here who's suffered the failures. We were in this space super early. We jumped all in in terms of adopting it internally, and I had some really hard conversations with my CEO about maybe being too aggressive. But he trusted me. He gave me the grace to roll forward instead of rolling back, and we have arrived at a place we never could have if he hadn't given me that trust and allowed me to fail but keep moving forward. So I agree completely with how you're breaking that down. On the failure side, what can go wrong? Take me through this, because I think people need to understand it and be ready for it.
I'm going to hit the first one, and we've talked about this a couple of times: unrealistic expectations. You don't just turn this on. This is not fire and forget, like, "All right, we're just going to turn it on and see what happens." I've been in IT for a long time. Sometimes we'd always say, "We're just going to wait till somebody complains and see what broke." That can be very, very dangerous here. I don't suggest that. Be curious about what could go wrong, but accept it. One of my favorites on this slide, though, is the legacy bias.
I agree.
If you're thinking agentic is just a different flavor of SOAR, you're way off. Way off. You've got to have your workflows and your processes, but if you're just doing a forklift over, you're missing it. You've got to rethink how you're doing it. You shouldn't be looking at reprogramming the agentic to do it only the way you think about it. What you should be doing is starting with: what's my signal? What's the outcome? What does success look like? But you shouldn't necessarily be defining every step in the middle. In fact, if you define every step, you're probably going to give yourself a bit of a crutch. Don't define every signal. Let the AI bring some of that to you. That's where you get the value. If you hamstring it so much, you're not going to win.
One of the other things: we talk about the criticality of human in the loop, and this is another area that could go wrong. Human in the loop is not the final stage. I know that could be a controversial conversation. Human in the loop is how you get to the final stage, which is autonomy. Because at the end of the day, if you're human in the loop, you're still human versus machine. The machines are attacking, you're fighting with the human, and the human's always going to be slower. You've got to get to machine versus machine. Build the trust with human in the loop. At the end state, get to autonomy. That's where you want to be.
I actually think that really harkens back to a previous slide we were talking about, which is the hype cycle and weeding through a lot of the vendor claims. If a company is just providing you agents, just a no-code agent builder or out-of-the-box agents, they're not going to get you there. There's a whole ecosystem of connectivity, workflows, marrying deterministic and non-deterministic to produce the type of precision we need. So I think your point there is spot on, and it really helps when filtering through the bizarre number of vendors in this space.
Yeah. And one other thing I'll mention real quick: if you get to a point where you are no longer asking questions, you're in a bad spot. With agentic AI, you should always be asking questions. Keep that curiosity going. Always question it.
BJ, I'm glad you brought that up, because we have a couple of questions coming in from the audience. One of them is: how do you keep token costs down as the agents scale up?
Tokenomics.
For BJ.
I'll hit this and then turn it over to you, Tim, because I'd love to get your perspective as well. Tokenomics is a key part of the conversation, and this goes back to balancing resources and execution. What I found to keep tokenomics down is being smart. This is where the human in the loop not only plays a good role in managing the output and limiting the impact; as you look at it, start to ask those curiosity questions. If you've got a runaway agentic, you're going to be spinning out tokens real fast. If you ask an agentic, "How did you get to that conclusion?" and it's parsing through data that's not relevant because it doesn't know the correct correlation, then you're going to be spending a whole bunch of extra tokens. So when you start to ask how or why it got to that conclusion, you start to look at the efficiency of it.
The other thing, and it's a little bit on the screen here, is be a tester. Go test your agentic. Test it not only for weaknesses and vulnerabilities, but for process execution. Test it for token consumption. And when you get to rinse and repeat, here's another thing: we talked about it being a crutch. If you always let agentic go do the same manual process to fix something and it consumes a large volume of tokens, maybe that's just a control you need to go block. Look at implementing a stronger control on a technology to prevent that situation. Sure, you can lean on the agent to always run and automatically fix it through its looping process, but maybe you just go remove a permission, change your access, implement stronger firewall rules or preventative controls. You've got to be curious about how it's executing and be smart about that. That's where you're looking at that output, because there are other ways to make it more efficient and save your costs.
Yeah. And I'll tell you, we had a no-code agent builder 18 months ago. The difference between what we had 18 months ago and what we have today is a massive ecosystem of tooling built almost exclusively around the ideas of delivering precision and tokenomics. We have all sorts of patterns in the product to enable an agent, for example, to spin up a script inside a sandbox to generate the JSON for you instead of using tokens, because if it does it with tokens you're going to spend a grand generating a JSON, whereas a simple script can do it for free. So there is in fact a ton that can be done to squeeze out 99%, and we have hard numbers on this, 99% of the tokens used in a purely agent-only system, versus a mature solution like Strike48 that combines workflows, scripting, and all these other technologies to keep the token lift small. And it's not a small topic. We could do a whole webinar on tokenomics alone. But at this point I can confidently say that no matter what problem a customer is trying to solve, we can do it at a low token cost. It just takes a lot of controls and a lot of different technologies.
Yep.
Richard, was there another question, or should we...
We do have a few more questions, but please keep going.
Okay. Stay on time here. I'm a little bit slow, as I always am. So BJ, if you could take us through the four-quadrant framework we've got here and the goal state.
I like this slide because we want to give you the ability to do a self-evaluation. Ask yourself where you're at. There are pros and cons to all of these. Look at the learner. Like you hear me say over and over: be inquisitive, ask questions, be curious. That's important. I hope that's where most people start. You've got to start with those questions. If you get to the skeptic, you should use that as a red flag. If you're getting to that skeptic, ask, okay, how much are you asking the questions? If you feel like you're not getting the ROI, you may not be jumping in deep enough. You may not be partnering enough. You may be looking too much at an out-of-the-box or black box solution. Dig a little deeper and you can transition out of that.
On the other side, some of you may get heavy expectations from executive leadership: we've got to move faster, faster, faster. That's something I hear all the time. "Great, can you have it done tomorrow?" Okay. Moving fast is good, but manage it with human in the loop. If you move too fast, you can trip, you can fall. Manage it with human in the loop, but you do need to get to automation. When you feel confident, transition over on the speed. Let it automate. Get it to autonomy. But I want to challenge you to get to the goal state of the tester. That means you've gone through it all. We've asked the questions. We had the skeptic going, "Ah, we weren't getting the ROI," so we changed things. Then we're moving fast and getting autonomy. But test it. Coming back to trust but validate: test it, validate the results, and look for the edge cases. Look for the corners. Look for the weaknesses in your agentic. It's great to red team it. Throw things at it. Do prompt injection, do poisoning, do whatever you can to try to trick it and maybe get it to do some damage you didn't anticipate, in a controlled way. You've got to get to that testing. If you're not in testing yet, that's okay, but keep that as your target and goal.
Yeah, all my bruises are from being perpetually in the speedster state. I've got to be more careful sometimes. That leads logically to speed versus trust: knowing when to release control, and almost more importantly, where to release control. Why don't you take us through how you bridge that gap?
Again, I'm going to start with the end in sight, which is you want to get to autonomy. So how do you get there? You've got to have trust in the use cases and the agents you're using, and you have to build trust with your constituents, your customers, the business members. Human in the loop is your transition. It's your way to do that validation: testing, questioning the hallucinations, adding more signals if you need to, bringing in more data, creating the correlations you need. That way you get to a point of validation. When you feel like you're doing checkbox validation, meaning all the output starts to hit what you're looking for, now you're at the point of, okay, we feel autonomous.
The hard part is jumping off that cliff and saying, "Okay, we're going to let the autonomy run." Don't be afraid to do that. Make sure you understand your critical points. Here's a suggestion, and we did this one ourselves: you talk about guardrails. You get to that point and say, "Okay, we're going to go autonomous, but here are the critical systems. We're going to leave this last 10% and keep some human in the loop there." Or maybe you restrict it, or do some additional monitoring. But let 90% of it run to autonomy. Then when you do that last bit of human in the loop at the end and you feel confident, you'll get to full autonomy. It's not black and white. It's a journey. It's a path. But you need to get to that target of autonomy. Don't lose sight of that goal.
Yeah. I love how you split that up and allow people to embrace autonomy in the less critical systems, at least out of the gate, because you're right: there is no success in an agentic adversary environment without autonomy. You're just waiting to get popped. All right. So how do we test our agents? How do we evaluate them like an adversary might, so we can get to that confidence level?
The testing is key. This is actually where I finally get to have a lot of fun, and it's also where I've had some great success with building trust. One of the things we've done in the past is go to the engineering group or the business, and engineers will say, "Hey, we know you guys are working in AI, so come let us beat it up." You give them the chance: all right, go break it. It's amazing, because you get a different perspective. You get people who aren't used to a cyber perspective really trying to push the edges. They're going to test it. They're going to ask questions. They're going to load things up and do all they can to try to break it. And that's the challenge: did you want them to do that? So I've had great success partnering with the business there.
We had one where we were doing a lot of email viewing and email responding with a lot of autonomy. People look at that as low-hanging fruit, but then you get white-on-white text with prompt injections that you don't want, and you get some weird results. Or putting out some honeypots, putting some data sources in the network, and the agentic stumbles upon it. Or in some of their config files they'll toss in something for the agentic. These are great opportunities. Another one: a lot of engineers like to get their hands dirty with the models. Give them a chance to beat up the full model. Let them see how it's being trained. They're going to try to put in different biases. It's a great chance to really open that up. Don't feel like you are the only one. Bring in a third party if you need to. That's always good validation for executives. I've brought in third parties to really beat up the entire AI platform and see what the results are. Last but not least, one of the best things when you get to testing is to know your controlled inputs and understand what you expect the output to be. That helps you figure out where the loopholes are. But don't be afraid to beat it up.
Yeah. I always tell people to take stock of who you have in your org, because particularly in this agentic world, some people are spending a lot of time upskilling themselves on these technologies, and you can have a real diamond in your org who can deliver you to a higher level in terms of testing your agents.
Very much so.
Okay. So what do you look for? What are the red flags and green flags in your AI solution?
Looking at green and red, the signals are always a quick and easy one. Understand what the agentic can consume. If you've got a limited set of signals, that's a red flag. You should have wide visibility. On the same note, if you're going back to the good old days of writing your own parsers on the log data, that's a big red flag at this point in time. You shouldn't have to write any parsers. It should be able to read it and see it come in: smooth, automated log ingestion. On the transition from human in the loop: if they're dependent on human in the loop, red flag. But if they're talking about the value of it, that's a good one. Pre-builts are nice to help you get started, but be questioning of those. If they've got nothing, that goes back to one of the biggest pitfalls of the SOAR and SIEM days, where you go to a vendor and they say, "Well, tell me what you want to look for." That means they haven't done enough of their own due diligence. So it's a big green flag to make sure they've got stuff coming through, but also to be able to separate out the agentic structure.
One I always look for: when you look at agentic, there's a big piece of the puzzle for identity. Is it RBAC? Is it their own independent accounts? My preference is delegated or inherited permissions based on the individual running it. It gives you a lot more control over the execution. It's a good guardrail. So make sure your agentic has flexibility in identity. Some of the startups out there are pushing for service accounts: "Give me a service account, I can't do anything else," and they're not integrating with the identity solutions. Identity is a huge risk in agentic. So make sure you've got a vendor and partner that fully supports a wide variety of your identity suite. It could be a major boon or a major vulnerability.
Yeah, that really rang true at Black Hat: the number of companies focused on identity, and then the number of agentic SOC players that had really ignored the concept in the effort to get a solution out quickly, was really stark, at least from my perspective. So, key takeaways. Take us through what you want IT leaders to know so they don't have to travel the same bumpy road you and I have traveled.
I'm going to hit them real quick. The anxiety is real, but don't be afraid. As CISOs, we often use FUD. I'm not a big believer in FUD, fear, uncertainty, and doubt. Don't be afraid of AI. Embrace it. Use it. Treat it like a workforce. It's going to need some teaching and training time. Help teach it to do what's right. If you just toss in an intern or an entry-level person and never coach them, they're going to cause a lot of damage. AI is the same way. Treat it like a workforce. Human in the loop is good. It's a key control for the journey. It's not the destination. Keep that in mind. And then beat it up. Test it all the time. I think it was Netflix that had what they called the chaos monkey. If you wanted to put something in production, it had to withstand the chaos agent, which would randomly shut down services and cause all sorts of chaos in your environment. You need to find a way to creatively test your AI all the time. Always try, always test it. Never trust that it's always going to be perfect.
But the most important one, at the very bottom, and this is the key one to end on: the time you put in is the time you're going to get out. Invest your time in design up front. Invest in the solution and it will come back to you in leaps and bounds. You'll get smarter token usage, smarter execution, better autonomy, better results, and more trust. These are key items. Put this together and you're on the road to success.
Yeah. And I will tell you, my experience has been that if you partner with the right vendor and put the time in up front, you actually end up with a much happier workforce. Not just happier because they don't have to spend their time on the wash, rinse, and repeat type jobs they used to, but because these technologies are really cool. I have seen SOC teams just light up when playing with our solution, saying, "Oh, I can accomplish this. Let me try to learn this so I can accomplish that." It's development of their careers. So it ends up being really positive.
People feel good when they feel like they made a difference. And I love what you said: give them a chance. They'll say, "Hey, look what I found," and it just builds morale. It's all-around awesome.
So listen, BJ, we ran a little bit over what I was hoping, but I think this was an incredible conversation. As always, thank you for your time. Anybody who wants a demo of Strike48 can scan the QR code here. And Richard, I'll hand it back to you if anyone from the audience asked anything you think we should address. Otherwise, I think we've gotten out what we wanted to and we can end it.
Yeah, I do have a couple more questions from the audience. The first one is: we have several years' worth of SOAR playbooks. Do we throw those out or have agents ingest them?
All right. I would say it's neither. You don't want to throw them out, but you don't want to bring them in as-is. Learn from them. Consider your playbooks guidebooks. They're guiding you on how you should be building out your agentic SOC. You've got a wealth of knowledge there. Not only do you have the playbooks, but you've probably used them, so you have the outputs. Leverage those and learn from them. It's a great jump start into that curiosity. It's a great jump start into validation. Use them. They're a gold mine of knowledge you'll want to take.
Yeah, and I'll second that. We actually have a whole ecosystem built around migrating existing workflows or playbooks almost as knowledge sets. Yes, for the critical ones, we can migrate them into an agentic system, but they're not agentic, so they end up as non-agentic in our system, there as a tool for agents to use. But what I generally find is in a well-planned system, if you're using them, as BJ was saying, as knowledge sets versus actual static playbooks, there's basically a 10 to 1 conversion. Ten playbooks translate into one agentic playbook that is much more flexible and handles the 10, plus probably 20 more, because it uses the 10 you had as a knowledge set to inform how to protect the environment. So I couldn't agree more with you, BJ.
All right, and we have time for just one last question. This one is: you said agents will fail sometimes. How do you manage that?
That goes to a balance with the guardrails, too. Here's my suggestion: if you've got a new, fresh agent you're working on, don't release it to your full production. It's a bad idea, because it will fail. It will hallucinate. And if you've got the ability for it to execute enabled, it's going to do things you don't want, and you're probably not going to want an output like that. So give it a controlled place. A sandbox is always a great place to start, looking at inputs, outputs, and success. Then slowly move it into your environment as you can. That's where you leverage inherited permissions versus giving it its own account. If you give it a god-mode account in your environment and just let it run, wow, it's going to cause a lot of damage. But with inherited permissions, it only has access to certain data, only has access to certain execution, only executes in certain places, so you're controlling the execution and the output.
That's the best way to do it, because it's going to fail. Anticipate that failure. At some point it's most likely going to fail catastrophically. Don't be afraid of that, but you've got to learn from it. Again, the more time you spend up front understanding how it works and how it reasons, the less likely you'll get a catastrophic failure. But at the end of the day, we need the automation there, the autonomy, because it is a machine versus machine fight. We need to get faster and faster at trusting and validating that it's doing what it needs to, and just let it go do what it needs to, because it's faster than any human is going to be.
Yeah. And failure doesn't have to be a scary concept. Catastrophic failure on critical systems, of course, is a scary concept. But if you're smart, you don't really need to face that. You can have failure in intent, failure in scope and comprehensiveness, versus catastrophic, "I just shut down a production server." So again, with continual iteration, staying curious, the tester ideal state, you can navigate from nothing, through failures, to autonomous. I've seen it done across our customer base. It's not even something that's elusive. It just takes time and dedicated effort.
And don't forget that failure can also be a false negative. Maybe it could miss something. So you've got to keep your eyes open there too. That's where you've got to know your outputs. If it didn't give you your output, maybe it missed something.
Although I will tell you, and this is no joke, we are at probably 20% of going in and actually finding active compromises in proofs of concept. So don't shy away from it just because there might be a false positive. The reality is a good chunk of the world is missing stuff as it is. So let's own that and embrace a better system.
All right. Well, it just remains for me to thank both of you for your time today. I learned a lot, as I'm sure a lot of people on the call did as well. So BJ, thank you, and Tim as well. This has been a great session.
Thank you, gentlemen.
Thank you very much. Bye now. Have a good day.