Posts

AI Factory Networking: A Simple Day-0 to Day-2 Checklist

Image
Planning compute for an AI cluster is more interesting work, and, accordingly, gets more planning attention. After that, network planning can get short shrift. As a result, there are many “gotchas” waiting to bit the unwary. This checklist attempts to address the planning concerns across the network for an AI cluster. Day-0: Planning This requirements specification describes how tenants will be isolated at the application and at the network layers. You probably won't need to purchase an additional stack because a lot of tools that integrate well with your platform (like enterprise Linux) already exist. Identify your requirements for support of disconnected or air-gapped environments. Day-1: Deployment Integrate your network setup with your workload orchestration to provide new tenants with network configuration automatically. You should turn on telemetry from the start of your project. Adding it later will be harder. Test that one tenant's traffic can't affect another's...

Bridging Workload Orchestration and Network Operations in AI Factories

Image
The transformation of enterprise infrastructure happens in two directions at the same time. In some cases, fast-changing and extremely distributed workloads need to be handled by enterprise infrastructures. On the other hand, enterprise infrastructures need to provide predictability and minimize operational costs. Managing multi-tenancy throughout the infrastructure and applications is vital in both cases. Also, constant visibility to application and network behaviors is required. With the increase in workload number and variety, quickly and efficiently resolving customer issues becomes increasingly difficult. Aviz Networks, together with Red Hat, provides solutions for workload and application lifecycle management with Red Hat OpenShift and Red Hat Enterprise Linux. For management of the network layer, they provide solutions at the Day-0 to Day-2 operations of AI fabrics. Aviz Service Node (ASN) provides a range of advanced capabilities such as deep packet inspection, packet-level dif...

AI's Hidden Bottleneck: The Network

Image
Networking for data centers is adapting to AI. Since GPU clusters are scaling up in size, they require very synchronized communication that can handle huge bandwidth as well as ultra-low latency. An Aviz Networks podcast features Taylor Allison, NVIDIA Senior Product Marketing Manager, who talks about the trend of networking with the current AI wave. During this conversation Allison touches upon how communication is essential in the AI training process, where GPUs transfer gradients back and forth, making network a crucial part of total workload performance. NVIDIA's Spectrum-X Ethernet, along with their InfiniBand options, has the capabilities to maintain performance in high demanding environments and help manage network related to AI. The interview further delves into NVIDIA's new Air feature that helps in deploying and configuring services by leveraging a Digital Twin feature, this can ultimately limit risk in planning for day 0 operations. Day 0 operations are by far the be...

What Happens Before a Spectrum-X AI Fabric Goes Live?

Image
The AI network doesn't just begin when the first GPU is plugged into a switch. A day zero in NVIDIA Spectrum-X involves topology definition, IP assignment, BGP setup, QoS settings, etc. At scale, the human effort that goes into configuring all of these components can turn into a substantial engineering effort. Intent-based automation can alleviate these difficulties. Starting with basic requirements such as Scale Unit and number of GPUs, IP ranges to utilize, Aviz ONES Fabric Designer can generate both the design of the fabric and configuration for the desired environment. This takes care of things on both ends of the system, from the IPCLOS underlay (leaf and spine setup, BGP configuration, ASN assignment, IP address, interfaces and MTU setup) all the way up to server-side setup (RoCEv2 parameters with PFC, and ECN, BlueField-3 settings, etc.). But setting up is only one hurdle. Validation is the other, so using NVIDIA AIR enables validation of the design, before rolling it out in...

Building High-Performance AI Fabrics with NVIDIA Spectrum-X

Image
Your AI network design needs more than just GPUs. Did you know that the Day 0 design decisions, topology, IP addressing, BGP, QoS, server connectivity, and more can have a dramatic impact on the performance of your network Fabric in production? Designing and deploying high-performance AI networking at scale using, for example, NVIDIA Spectrum-X is challenging. You can simplify Day 0 design and deployment with Aviz ONES Fabric Designer. Instead of hand-configure complex IPCLOS Fabrics and dealing with hundreds or even thousands of parameters, the ONES Fabric Designer lets teams input data at a high level – specifying the Scalable Unit (SU) size, the desired number of GPUs, and the IP address pool. The system then generates a tested and validated network Fabric design including spine and leaf configuration, BGP peering, interface settings, and IP assignments. Furthermore, the solution also fully automates the setup and configuration of NVIDIA BlueField and ConnectX SuperNICs including a...

Why AI's Biggest Opportunity Is Moving Beyond Hardware

Until recently, companies sold innovative hardware. Now companies are focused more on integrated and proprietary software. AI is disrupting companies' old ways. Quickly upgradable chips and networks are no longer the most important hardware to own. Long term value comes from data, models, and workflows that utilize AI to run applications. The biggest challenge to AI businesses is: Fast-developing hardware generations. Hard-to-keep-up-with upgrades in infrastructure. Dependence on hardware for developing applications. Growing demand for deploying AI quicker at a larger scale. The answer is not decreasing hardware spending. The solution is spending to develop hardware that is more flexible and evolving. The adaptability of infrastructure allows organizations to upgrade hardware without the difficulty of rebuilding applications. Integrating upgrades in hardware allows organizations to keep their AI workflows running and evolving. Implementing operating standards is what the market is ...

Can Today's Networks Handle Tomorrow's AI?

The way data center networks operate is changing with the evolving nature of AI workloads. In a recent Aviz Networks podcast, Taylor Allison, Senior Product Marketing Manager at NVIDIA, spoke about changes that will happen in networking when it comes to large-scale AI training and inference. To do AI training, GPUs need to work in tandem, which requires a network that is able to communicate with low latency and stay synchronized. Once you increase the number of GPUs, even more, the existing networking technologies will start to fail. The networking solutions from NVIDIA that are optimized for AI workloads include Spectrum-X built for AI clusters over Ethernet and InfiniBand, another high-performance networking architecture for AI overclusters. The podcast also covered NVIDIA Air and the digital twin technology. Before deploying configurations and automation to the production environments, users can create virtual environments in which they can test configurations and automation. This g...