How Do You Keep a Network Reliable When an AI Factory Scales to 10,000 Nodes?


Scaling AI systems mainly focuses on computational units. Networks are considered "plumbing" and assumed to work. In reality, there are significant differences in architecture and behavior of AI systems between hundreds of nodes and tens of thousands of nodes. Unfortunately, many AI projects remain "stuck" at the factory stage due to the inability to scale beyond hundreds of nodes.

Why does scale change the problem instead of just making it bigger?

Nicolas Kremer, CTO and Co-Founder of Polarise, put it plainly on a recent episode of The Human Voice in the Age of AI:

"I have this great solution. But then I ask them, okay, how do I scale that across 10,000 different nodes? And then it gets interesting, right?"

The interesting part is not just a bigger version of the same problem. The problem is multi-tenancy, predictable automation, and failure recovery all need to work together to support the main functions of the system. This is typically not an issue during a proof of concept study.

What actually keeps the network reliable at that scale?

Vishal Shukla of Aviz Networks describes predictable automation as the solution for managing multi-tenancy at large scale. Shukla believes that as networks increase in complexity, they will require dimensions of growth beyond the vertical and horizontal. Within a production facility, multi-tenancy can be used to logically partition a network to provide different levels of SLAs. Predictable automation provides this framework.

His three-part framework for doing this without losing control:

  1. Process — Your reference architecture should be set before you start implementation.

  2. People — the correct experts who can determine where risk truly begins, as opposed to who is most convenient.

  3. A predictive platform — the automation you build from what the first two pieces teach you, not a generic tool bolted on afterward.

Why can't one company just build the whole stack themselves?

They pointed out: no company can solve AI factory scaling on their own. In order to solve these problems at large scale and with greater quality, NeoCloud must rely on partners, as most organizations are attempting to develop in-house solutions for compute, storage and networking and build out their own datacenters.

vCluster's Lukas Gentele discussed the automation needed at the software level, and stated that beyond a proof-of-concept or demo, it is essential. An example he gave was a power outage causing a grid to go down, leading to a data center outage. This offers a scenario to put an architecture stress test. In the scenario given, if the architectural design relies on a person going to a data center and manually running a script via SSH, then the design and architecture would likely fail.

The part most teams underestimate

In closing, David Iles from NVIDIA told the audience to not to treat their data centers as science projects and, instead, stick to time-tested best practices. Data center architects can get captive to the latest trends and consider each of their deployments as new and different. Many of the ways that deployments can fail are fairly consistent. During the presentation, the panel touched upon some interesting paradoxes within the storage industry. One of which is that for storage, multi-tenancy means different things to different companies and that differences, in turn, make up one of the most difficult issues the storage industry needs to solve.

If this was useful, follow for more on AI infrastructure and open networking. Want to go deeper on making AI networks production-ready? Visit [aviznetworks.com].

Comments

Popular posts from this blog

Scaling Deep Network Observability for 5G: Reflections from a Real Deployment

Evolving Packet Brokering for Modern Network Observability

How Network Copilot Uses Agentic AI to Correlate FortiGate and Splunk