LIVE · HOW WOULD YOU LIKE TO CONSUME THIS PAGE?
Close
Privacy settings
We use cookies and similar technologies that are necessary to run the website. Additional cookies are only used with your consent. You can consent to our use of cookies by clicking on Agree. For more information on which data is collected and how it is shared with our partners please read our privacy and cookie policy: Cookie policy, Privacy policy
We use cookies to access, analyse and store information such as the characteristics of your device as well as certain personal data (IP addresses, navigation usage, geolocation data or unique identifiers). The processing of your data serves various purposes: Analytics cookies allow us to analyse our performance to offer you a better online experience and evaluate the efficiency of our campaigns. Personalisation cookies give you access to a customised experience of our website with usage-based offers and support. Finally, Advertising cookies are placed by third-party companies processing your data to create audiences lists to deliver targeted ads on social media and the internet. You may freely give, refuse or withdraw your consent at any time using the link provided at the bottom of each page.
Thank you! Your submission has been received!
Oops! Something went wrong while submitting the form.
WEBINAR

PT

PT

Demo: Observe What Your AI Is Actually Doing

Security teams are receiving more alerts tied to AI workloads, but most miss the runtime context needed to understand what happened, why it happened, and whether it violated policy. AI visibility cannot stop at deployment and configuration.

Join this live demo session to see how Wallarm AI Hypervisor helps teams understand what AI workloads are actually doing at runtime inside Kubernetes environments. The session focuses on giving security teams clearer operational context around AI behavior, outbound activity, sensitive data exposure, and user-driven actions across AI systems.

In this session, you’ll learn:

  • How AI Hypervisor instruments workloads without application changes
  • How runtime AI activity is mapped to users, sessions, and outbound connections
  • How sensitive data exposure is identified across AI pipelines
  • Live answers to your AI runtime security and governance questions from Wallarm experts


Thanks for the download. 
Thank you for registering for our webinar, you will receive a confirmation email shortly. We look forward to seeing you soon!
Outlook
Google Calendar
Apple Ical
Outlook.com
Thanks for filling out the form!
The webinar link will open in the new tab. If its not, please follow
this link

Our Speakers

Ash Patel
Sr. Partner Sales Engineer, Wallarm

Ash Patel: Thank you, Annette. My name is Ash Patel. I am a partner solution engineer here at WallArm. And I'm excited to tell you about a new product we've been working on for a while called AI Hypervisor, or — I like to think of it more so as our AI control plane. Before we jump into the demo, I want to set the stage for why we're here.

We're seeing a massive shift over the last year. AI has moved from the experimentation phase and is now a core part of many business workflows. It's not just a side project — it is core infrastructure for most organizations.

However, we're also seeing a significant gap. While adoption is accelerating, the flip side — governance — has not kept pace. McKinsey research shows that 88% of organizations are using AI. If you're not using AI, you're probably going to be left behind. But only 30% of those surveyed said they've reached any meaningful maturity as far as their governance around strategy and controls actually goes.

This has created tension between security leaders — the people tasked with securing everything — and the people asking the questions: the auditors, the CEO, and others who just want to make sure things are being properly safeguarded. 80% of people surveyed in a recent industry survey said they've already experienced some type of incident involving generative AI. That's broad — it could be detecting somebody sending proprietary information to public LLMs, or an LLM giving PII to somebody who shouldn't have it. So that can come in many different forms.

Ask yourselves how this disconnect between adoption and lagging governance is playing out in your environment. We probably have a lot of security practitioners and security leaders here — are you seeing this friction? The answer is probably determined by where you are in this journey. AI transformation is crawl, walk, run — there's definitely a stepping stone to good governance.

But we see most organizations we talk to following a very similar path. It starts out with AI as an experiment — call it the discovery phase, where a team, or even an individual who likes tinkering with AI, builds something internally and says, "look at this, isn't this cool?" It's uncoordinated, which is probably expected at this experiment phase.

If those experiments show promise, get some eyeballs, and get a sponsor, they begin to scale. This is when somebody with budget authority and political capital — maybe even the CEO — says, "we're all in, we're going AI first." Then you see an explosion of projects and usage, and the primary challenge becomes keeping track of what's going on in the environment.

If some of those projects become successful, teams move on to operationalizing — this is where AI moves out of a lab/testing phase into being part of an actual production pipeline. And finally, there's the governing side of it — once you've operationalized it, the goal is to implement controls to manage AI at scale.

But the question still remains: what does good governance actually look like? Are people still tracking everything in spreadsheets? Is there a bunch of disparate tools where they're trying to stitch together — "well, we have this in a CMDB, and this in a log management platform, or this in an observability tool"? We hear all kinds of answers to that.

That's really the reason we're here. When we talk to teams putting AI into production, we find that the conversation usually narrows down to four critical functions, which map to the technical challenges of maintaining control:

First is inventory — what is actually running in my environment? Second is behavior — what is all that stuff actually doing, and what can it access? Third is control — am I looking for malicious activity, and if I find something, do I actually have the ability to control it, or to detect risky behavior before it becomes an issue? And the last piece is evidence — if an auditor comes in, can you actually prove that this is under control? If somebody from GRC, who cares a lot about DLP and data loss, comes and says, "I need you to show me — not just tell me — that you have this under control, and that no PII is getting out, and we have guardrails in place" — you need to be able to show that if there's any type of bad actor, hallucination, or rogue AI, it's not going to have access to the crown jewels and be able to break out of a sandbox.

And so this is where WallArm comes in. We start out by helping you discover — and just for clarity right now, this is an AWS-specific platform; we've partnered with AWS on the launch. We're targeting workloads running in EKS right now, but there are plans this year to broaden the scope to the other major cloud providers.

So getting back to it: first is discover — knowing what is running. Second is observe — knowing what the AI is actually doing, understanding how AI interacts with users and how users interact with AI, what types of data the LLMs should be able to access, and whether there are defined boundaries you're able to show. Third is enforcement — if you have a user who knows he's on the way out and decides to scrape or harvest as much data as he can, and starts trying to jailbreak prompts by telling the LLM he's a super admin who should have access to everything, that's the type of activity you'd want to know about and be able to control the response to. And the last portion is proving that the AI is under control — providing compliance evidence and having a thorough trail of AI risk reporting in place.

So, I did not want to do death by PowerPoint, but I needed to take a few slides to set the stage for what I'm going to show in the platform. As you can see, there's a lot in the platform, and one of the underlying foundations we set out to implement in the product from the start was to have different views.

Because AI is touched by so many personas in organizations — from executives to security engineers, compliance people, down to developers — everybody has different needs, and everybody wants to see different things in the platform, so we try to cater the views to them. I don't think I'll have time to get into all the views today, but I'd say the two most valuable are the executive view and the security engineer view, which we'll look at in just a moment.

If you're an executive, if you're a CISO, you care about high-level stuff — you want one pane of glass where you can log in and it'll show you what is running, how many AI apps you have, whether you're watching and actually blocking anything, whether there were any potential PII incidents, what the user adoption looks like (is this all agents, or is this actual users), and then a prioritized view of things that need decisions.

So, we try to make that very easy for the executive. Supply chain is a big thing right now — we hear of all kinds of NPM packages being compromised, people deploying them without real verification. You're probably guilty of that at home; if you're doing it in a large production environment, the risks are a lot greater. So we'll show you the CVEs we've currently found and other types of vulnerabilities — things you should be aware of that you can then delegate to others to prioritize.

On ungoverned AI: when we deploy, there's a configuration process — you tell us which namespaces we should be watching closely. But just because we're watching those namespaces doesn't mean we can't see everything. We're running at a very low level, using eBPF along with a combination of proprietary, patented memory scraping, to really see what all the containers and pods within that cluster are doing. So we can say, "you have 50 namespaces here, yet only 45 of them are being monitored by AI Hypervisor" — that might be completely intentional, or it might not be, but we're going to tell you either way. From doing that, we can also show you if you haven't sanctioned a particular LLM provider — say you wanted to lock your environment down to only use OpenAI, but we see people using DeepSeek, Grok, and other LLMs outside the scope of what you're regulating. There might be ungoverned AI spend, which we'll get to in a moment. And you might have activity that has no behavior certification — which is where I want to pivot for a moment to talk about certification, because the A2AS framework is one of the foundational philosophies of our platform.

So what is the A2AS framework? Essentially, it's a framework to govern AI agent security and monitor AI agent governance. You can see here there are a lot of big names involved, and then there's WallArm — we're not as big as the rest of these companies, but we try hard, and we've shown a lot of thought leadership in this space to be coordinating this effort with a lot of big-name companies.

At its core, think of the A2AS framework as a declarative sandbox of what agents can do. At its most basic structure, it's a YAML file — but within that YAML file you can set controls around authenticated prompts, security boundaries (is this LLM able to talk to just MySQL and not Postgres, for example — that's something you can implement), in-context defenses, and codified policies. There's a lot in here, but just know that this framework, and the certification of the traffic you're seeing, is essentially the foundation for AI governance — and that's the type of thing we'll help you build.

So when we say "no behavior cert," it means we're seeing traffic patterns — it could be an MCP agent, it could be a specific user, whatever the case — and we'll show you if we have traffic that is certified: we've been watching the traffic for a while, we feel comfortable with the boundaries this LLM is able to reach, and once you've seen enough data, you have the ability to go into the platform and actually certify or revoke the communication if you're planning updates.

We can see here our behavioral certs — authenticated prompts, detected traffic served over TLS (which is what we actually want), and security boundaries — there's Nemo, the NVIDIA LLM guardrails, in place. I won't go into the nitty-gritty here since we're still talking about the executive view, but the executive who logs in can see if any agentic or MCP-server-based communication is or is not certified, and based on that, delegate to people — "we have these five mission-critical applications, I need to get these certified as soon as possible."

While we're here — this is the registry page. You'll see these chiclets at the bottom; the registry chiclet gives you a good overall picture of everything running in the environment. Here we can see the agents that are running, the one single MCP server, the LLMs being called by the various applications and agents and traffic, and the APIs — we love APIs here, and these are effectively the front door to all the agentic communication, so we can show you where the source communication is actually starting. I don't think I have any data connections here — we just have one Redis data store in this test environment — but if you had S3 buckets or other types of databases, they'd be listed here in the data section. And to round it out, these are the certificates we were just looking at a moment ago.

Now, one thing we found is that executives care a whole lot about spend. This is one of the foundational pieces here too — as a company, I can assure you WallArm is spending way more than 14 cents on our AI, but we look at token pricing for various LLM providers out there, so we can give you a good picture of where your spend is actually going and break it down by users, by agents, by role, etc. — however you want to slice it.

Okay, I think that's enough time on the executive view — let's flip to the security engineer view. But before we get there, I want to show a couple of things related down to the application level. We saw the registry component, but if you want a more visual topology map of how the inter-cluster communication is actually happening, we can show you that here — a clearer picture of egress. We don't have egress coming out of this test cluster I built, but if we did, we'd see it here. Then there's the data governor, and lastly, enforcement.

You want to have rules in place — this is what the governance people actually care about. You can have rules related to prompt security and data protection — so if you see potential PII exfiltration, you can activate a built-in rule (you don't have to create it from scratch). Within the various enforcement components in the product, you can have rules to prohibit jailbreak attempts, prevent prompt injection, or trigger scanner flag session alerts. Other runtime controls include high-risk session detection — if we see a behavioral pattern (we'll look at behavioral patterns in a moment) that looks a little fishy, you can just block it off the bat — say, "if this type of activity is high or critical, go ahead and block it."

And then here is the MCP server tools. People stand up MCP servers all the time without permission, just because they have the ability to, and no one's really watching what capabilities those MCP servers and tools have. So if we see rogue MCP servers in the environment, you have pretty granular actions you can take — block, alert, log, redact. If we see PII going out, we can actually capture that, because we're sitting as a proxy — we can take it and redact it, putting in a bunch of asterisks before it ever reaches the end user. We'll see that in just a moment here.

If you're a security engineer coming in, chances are you're going to spend most of your time looking at things like this. You'll come in and say, "show me the patterns I'm seeing here." All the data and telemetry we're bringing into the platform is being fed into an LLM itself to help us look for patterns — and we can see here that PII reached an external LLM in 30 separate sessions. If we drill down a little deeper, we can see a summary — this is the redaction, so we've taken that information out. If we click down into the session — I think this is because I'm currently read-only, so it's showing that — but if you click down a bit deeper, you'd actually be able to see something specific down to the prompt.

So let's go in and look at this session — we can see here, "[email protected]" (this is a fictional email address, a Matrix reference). Having the ability to look at potential PII exfiltration, drill down to a session level, and then actually implement rules around blocking it — because I'm read-only in this internal environment, I can't manually kill the session from here, but in real time, if you discover these threats, you have the ability to block that request or kill that particular session.

What this means here is that currently, in this environment, we have 17 active sessions — they could be agents, they could be real users. "Tainted" means sessions we're seeing that have some type of PII or other sensitive data that should be flagged for further follow-up. We can see here "[email protected]" — hopefully everybody's watched Silicon Valley and knows what Hooli.com is — but this is the smoking gun: when somebody says "no, I didn't set up a rogue MCP server," you have the evidence right here. You can see that [email protected] — this is his session, this is what he was actually trying to do.

And so this is the audit trail — this is the governance side of it, where the GRC people come in and say, "I like these rules, I like that you're able to identify social security numbers, physical addresses, emails, and things that could get us fined if they got out." They like knowing they have the ability, from the session interface, to block that PII ingress and to implement controls in the form of the behavioral certificates for all the AI communication in their environment.

So, from a high level, that is the WallArm AI Hypervisor platform. I don't like doing river cruises — I don't enjoy giving extraneous bullet points, which I think get us away from the core message of the story. So, Annette, at this point I'll wrap it up on my side and open it up to see if there's any questions.

Annette Reed: Awesome, thank you, Ash. Alright, feel free to raise your hand if you want to ask a live question, or drop a question into the Q&A or chat. We will stay on for a couple minutes to answer any questions that you might have.

Ash, have you seen anything — any common questions that you've had so far?

Ash Patel: I've gotten a lot of questions, but some of the common ones are, "how long should I observe data before I'm ready to certify it?" We kind of default to seven days, but it really depends on the amount of traffic and events we're seeing. That's one of the big questions.

Annette Reed: Alright, it looks like we have a question that's come in through the chat. Do you want to—

Ash Patel: Yeah — is there anything to identify the provenance of an AI agent, like agent IDs? Let me see here. So, there's an internal ID, but I think I need a little more clarification — what do you mean by the provenance, like who deployed it?

Annette Reed: Yeah, let me find — yep, let me get you off mute. Sorry, let me just find your — where did you go? There you are. Okay. Alright, you should be able to come off mute and ask your question.

Sai Ramanath: Awesome. Can you guys hear me?

Annette Reed: Yes, we can.

Sai Ramanath: Awesome. Thank you, Ash and Annette. So, the question is about — like, which agent was used, which provider — is there any traceability toward that? Because of the recent OpenAI/Hugging Face thing, there are some governance considerations from Europe which say that every agent must have a unique ID and that kind of thing, so I was wondering if WallArm provides anything right now — does the agent emit any such thing? Do you have any provision for that?

Ash Patel: I would say not today. But that's an interesting use case, especially knowing that the AI controls in Europe are way more stringent than they are here — I know they recently introduced some AI regulation, so that's a point I'll take back to our product team.

Annette Reed: Yeah, and we actually just — great question, and perfect on the timing, because we actually just published a blog yesterday, and promoted it on our social today, around that incident, so you might be able to read a bit further on that. I got a bit more information that came from one of our lead DevOps engineers, so it should be a good digest of what's going on, and you might be able to get more information on that capability.

Sai Ramanath: Annette, I somehow cannot get enough of this one, so this will be great. Thank you, Ash. And again, also for open-weight models that you self-host, etc. — there must be some way to connect the dots, like who used what, that kind of thing, so that would be very useful. Awesome, thank you guys.

Ash Patel: Yeah, no — so if we go into here, a lot of this — even the application I'm running here — a lot of it is on OpenRouter on the back end, so we're calling all kinds of different LLMs, and we saw a couple moments ago with the "dsomething" example that we can see exactly who is prompting what. Not just who, but we can see the exact prompt they're trying to run.

Sai Ramanath: Thank you.

Annette Reed: Great question. Any other questions? I have not seen any new ones in the Q&A or the chat, so feel free to — you will get an email tomorrow with the recording of this webinar, as well as a couple of the resources we shared today. And of course, if you want to go deeper into this, you'll also see a "Request Demo" link, where you can get a one-on-one demo with someone like Ash, to walk through a lot of the scenarios, very personalized for you.

So, thank you again for joining today. If you have any questions I will stay on for just a couple more minutes, but we'll go ahead and close. Thank you, everyone.

Ash Patel: Thanks, everyone.

Our Speakers

Ash Patel
Sr. Partner Sales Engineer, Wallarm
Trusted By

The world's most demanding teams run on Wallarm.