Talk title: “AI Infrastructure Challenges — Why Intelligent Network Operations Matter”
Speaker: Ronny Wolf, Field CTO, BlueCat Networks
Event: Infotech Barcelona
Duration: 19:54 ·
Language: English ·
Machine transcription, lightly cleaned for punctuation. Uncertain words are marked with [?].
[00:00] Today I’m going to be here at Infotech Barcelona. I hope you have a successful day — already a successful day. My name is Ronny Wolf, and I’m from BlueCat Networks, Field CTO. I’m wondering, what is a Field CTO? I’m a little bit the interface between our [CTO office?] and our customers. So, sitting at customers, listening to them, and sometimes also thinking about how we can solve things there.
[00:34] My topic today is AI infrastructure challenges — why intelligent network operations matter. And with that, I will slightly start.
[00:50] For sure, you have seen in many presentations today already these kinds of “whatever AI is,” digital transformation, blah blah blah, things like that, right? But if you think about that — I mean, over the last decades, we already transformed applications: virtualization, outsourcing, insourcing, cloud. Then we started with automation. And maybe we are still in the transformation process, right? I mean, it’s an ongoing transformation process — whatever, from cloud back to the data center and vice versa, and so on.
[01:25] And now the big trend and buzzword, AI, comes into play. It’s the next transformation, which basically affects all of our businesses. But the interesting thing — and if you also take a look at the booths, if you take a look at whatever, at all of these marketing slides: AI is easy to use, everybody can adopt that easily, and stuff like that. But one thing that most companies really don’t have on their radar is that with AI, your network is under more pressure than ever.
[02:03] Every AI interaction depends on the network, right? That means DNS lookups, service discovery, additional API calls. We need to make sure that the LLM — maybe it’s on AWS Bedrock, or it’s on Azure AI Foundry, or something like that — that the connectivity is there. That our AI system can communicate. That we have also the proper security policies in place. And that the AI models can really deliver what we want.
[02:27] So the question basically is no longer whether your organization will adopt AI. It’s whether your network operations are ready to support it — to really be successful with AI.
[02:49] A couple of examples. I mean, everybody knows, I think, we are talking about copilots — whatever, Anthropic copilot, Copilot from Microsoft, and so on, OpenAI, similar things. So basically, our customers or our users are already using that within their business workflows. We are using that in software development — Codex, Claude Code, and so on — to be a junior developer, or maybe even a little bit more, a software architect and stuff like that. We are using it in IT help desk, customer services. We are using AI agents across the business.
[03:35] And the next step, for sure, is then also — even if you’re not at that stage already — that you will let AI do things autonomously. So with autonomous agents, they automatically do certain stuff within your network or within your applications. That means we have to deal with AI models, whether it’s on-prem, cloud, as a service. We have front-end, we have back-end, MCP servers all over the place, AI agents which are connected to it.
[04:05] So what happens when every employee and application in the business starts talking to AI, right? And again, I mention AI as a big overarching term — you have seen whatever, these copilots, MCP servers and all of these words.
[04:25] In traditional networks, it was quite easy to predict. So without AI, and without that complexity around it, it was quite easy to know: okay, in the morning everybody enters the building at 9 a.m. with [unclear], everybody is using whatever, Office 365. Hey, we know north-south, east-west traffic, right? So we can predict that. We can scale our network accordingly. Quite easy.
[04:55] With AI traffic, it’s not that easy anymore to predict. And don’t think about — you have your ChatGPT, right, somewhere, you talk directly to OpenAI servers, they have the LLM in place and things like that. That is an easy thing, right? But this is not where we are successful with AI. Our applications talk AI. Our users talk AI. Maybe they autonomously talk AI to each other, with MCP servers, with MCP clients, with AI agents. Then maybe some automations in the back end are running some stuff, and so on. So it gets even more complicated.
[05:40] And I’m not even talking about hosting your own LLMs, neural networks and stuff like that in the back. That gets even more complicated, where you then just scale that even more.
[05:47] I mean, that is a report from Logicalis. They did it last year and released it beginning of this year. AI adoption is slowed by major challenges, and 80% said, okay, our infrastructure limits our ability to scale AI really to what we need.
[06:00] And why is the network a bottleneck in that case? I already mentioned one thing, right? Complexity. It gets even more complex. Our network operations guys don’t even know who is talking to each other. That was already complex when I said cloud transformation, virtualization and stuff like that, automation — it already was complex, right? But it gets even more complex, and we don’t have that predictability anymore. There is potentially not even someone we can ask: “Hey, cloud team, do you know what sits there?” I have no clue where they’re talking to each other anymore. Visibility is an issue.
[06:38] So from complexity we get the complexity, and then potentially also the latency, right? If we introduce latency within our network, an end user potentially complains: “Hey, my internet is slow, my business application is slow,” things like that. But if we have autonomous agents who should do some work automatically, or maybe during the night — whatever, do a couple of GitLab commits or things like that, doing automatic testing and stuff like that (again, AI testing I mean here) — if your latency comes into play, you start in the morning and come to the building and every fifth [?] test fails, whatever, the automation cannot proceed, and so on.
[07:29] Yeah, and that comes to the conclusion — and also we have hybrid infrastructures, potentially security policies in place. We have more dynamic workflows than ever. And for sure, we depend on a really stable and resilient core network, or core services in the back end. So most AI failures are not AI failures — they are infrastructure failures.
[07:56] And one of the things traditional network operations have a challenge with: they are siloed. And they are siloed for a reason, right? I mean, the security guys need to sit somewhere, they need to have their say and say, okay, you cannot do that, you should not use whatever — based on our governance, based on our standard policies — to use a certain AI tool or certain communication and stuff like that. We have the cloud guys: they are agile, they are fully automated, potentially with Terraform, Ansible and stuff like that. We have our network engineering guys who are architecting the network. And then we have the network operations guys who are looking at whatever, there’s a red bubble, something is not working properly, maybe our routers, switches, things like that have some latency issues.
[08:43] And all of that has basically to make sure that they interoperate in an AI world seamlessly. Doesn’t matter if it’s the core network services like DNS, DHCP, NTP and so on, application logs, access rights, maybe the cloud infrastructure — and here’s also only an example of cloud logs and stuff like that. Our whole network infrastructure: routers, switches, access points, WAN connectivity, virtual machines, containers, and so on. And then we have the big security guys who need to make sure that firewalls, proxies, endpoint detection and response systems and so on are working properly.
[09:34] And as you potentially already recognize, maybe in your own organization: the network operations guys see something is not working. They need to open a ticket, whatever, to the cloud team or to the security guys, and it takes a couple of — some time, and so on. But think about that we are in a fully automated world at the end. We want to be faster with AI. We want to be productive with AI. That harms us in this traditional network operations model.
[09:59] So the idea with intelligent network operations is to make sure that all of these things come together into one single platform. So the cloud teams, the security guys, and so on — so that basically their shared goal is to have an always-on network, with all the capabilities we need to keep up that pace, to make sure that these requirements are fulfilled.
[10:33] And the idea behind the intelligent network operations model is that we operate across all the critical control and visibility points. That means one single platform. We need to understand where our whole underlying network infrastructure is: routers, switches, [antennas?], servers, things like that, but potentially even physical stuff and so on. Then on top of that, we need to make sure that we have the core network services provided properly, right? We need to get an IP address. We need to make sure that connectivity is working.
[11:09] And guess what? AI is highly dependent on DNS traffic. So we also own that and have visibility into the DNS service itself. And that we know where all of these devices are, with a proper IP address management.
[11:23] And if we have all of these control points and visibility points, then we can also have a platform which is really then bringing all of that stuff together — from a security perspective and from an observability perspective. And if we have that stuff at the end, then we can also think about automating even faster, to bring all of that stuff together. New router, new switch — automatically it will register via DHCP into the IP address management. It automatically gets the alerting or the monitoring stuff. We automatically capture packets, flows and stuff like that into it.
[11:59] And we also have the ability not just to gain the visibility, but also to have control around it. And if we have the control around that, then we are very close to what we say here: an intelligent and self-healing network. If we see there is something not working, automatically adjust the [driver?] again on the router, on the switch, maybe provision something in addition, expand when we get out of IP addresses, automatically make the network larger and things like that.
[12:41] And if you provide real-time visibility into traffic, performance and behavior, then you also can think about — okay, network operations is not only about that we see all of that stuff, and you potentially will say, okay, this is a network monitoring solution, we have that already in place, right? We have our observability platform. But what if we change something in the network? Does it really understand the intent of the change and the risk associated? Do something on your core network services, whatever — remove a network, expand that network. What is really then the impact if I change something like that? Is it high risk, and can I delete that safely? Is it — for example, on a security perspective, change some firewall rules: what is affected by that? Oh, your Claude Code cannot communicate anymore with an MCP server on the other side.
[13:45] And then you can rate the risk. We bring that together, you see that in one single platform and say, “Oh, maybe we do that change not now,” otherwise something is falling apart. So intelligent network operations basically understands what’s happening, understands why, and will help to drive your actions and reduce risk.
[14:10] Guess what BlueCat is also leveraging? Not just that we have the ability to make sure that your network keeps up that pace in an AI and this transformation — to a more AI-driven network — we also leverage machine learning and AI within our platform itself.
[14:34] Because if you think about that — and maybe I go one slide back — with all of these routers, switches, [antennas?] and so on, you have your platform over there, for sure. Whatever, Cisco Catalyst Center or something like that. Or maybe you have a network monitoring system in place there. Then you have something for your core network services — doesn’t matter if it’s Microsoft DNS, DHCP, [?]. Maybe you have an open-source IP address management. You have a little bit on NTP and all of that stuff. Then you have your observability platform in place, and so on.
[15:12] And if you think then — all of that stuff is producing a vast majority of data, right? It’s a lot of data: flow data, packet captures, security policies, firewall rules, whatever it is, network device configurations, and so on and so forth.
[15:25] So the idea basically is that we leverage AI with an autonomous agent as well as with an interactive agent, that you can ask: “Okay, what is the impact if I do that?” And not just looking into whatever, a nice diagram and a couple of packet captures and flows and stuff like that, and then someone sits behind it and says, “Okay, if I do that, then I need to do that, and then I go there, there’s a drill-down here. Okay, is that application affected?” You just ask the AI — it’s connected to the platform itself, and it tells you exactly: okay, this is what happens if you do that change.
[16:16] For sure, that is an idea — yes, we have that in the product itself. We know that you will not rip everything off, right? And so the importance with that is you need to integrate it. And that is a complex picture, I know, because I created that with AI as well. So whatever, you see that typically, right, that AI creates these kinds of things. But it was the easiest way [?] that I found.
[16:42] So while you potentially have the platform of BlueCat in place, we know that you have your own applications, you have your platforms that you’re managing right now, you have your automation platforms, your whatever. So the idea basically is that you can connect all of them together.
[17:08] Typically we are doing that with MCP servers. So the platform itself provides MCP servers. Our agent is open, so that you can connect your existing MCP servers. Or maybe you have some kind of standards already defined, right? You have already decided we are using Anthropic, or we are using OpenAI, or we are using something locally — so that you can connect all of that together.
[17:35] We even look at the platform from all of these core services: log servers, packet capture service, inventory config service and stuff like that, bringing that together. Potentially the answer is a little bit different than using our AI agent, right? Because we are using some kind of guardrails in between, and so on. Yes, the MCP servers also have guardrails, but at the end the result would be very similar.
[18:00] And then you can get really from your clients towards the application itself — it’s now here in the internet, or wherever it is. And all of these different data points will be collected automatically, again, because these MCP servers are provided to the AI agent.
[18:26] And whether you’re doing that automatically — some stuff, in guardrails again — also our, what we are calling, LiveAssist, can automatically do certain stuff. Maybe not delete a network, maybe not rip out a router or switch. So again, all in guardrails, and you have for sure your own policies in place later on. But then it understands the intent based on that information. It correlates all of these data together and provides you, in the first place, with more visibility than ever, more intent than ever — with the change here and something happens on the other side. And yeah, integrate them then into your existing ecosystem: the monitoring systems, SIEM, SOC, ITSM, ticketing systems, and so on and so forth — to leverage all of that data, to leverage all of that knowledge around that, to keep up the pace on your network.
[19:16] With that, I think I am already over my time. With that being said: if you want to see that in action, how we are doing it, how you can connect this to your existing network infrastructure, if you want to learn more about BlueCat and our platform approach — visit us here at our booth. Thank you very much, and enjoy the rest of the day.