Posts

Bridging Workload Orchestration and Network Operations in AI Factories

Image
The transformation of enterprise infrastructure happens in two directions at the same time. In some cases, fast-changing and extremely distributed workloads need to be handled by enterprise infrastructures. On the other hand, enterprise infrastructures need to provide predictability and minimize operational costs. Managing multi-tenancy throughout the infrastructure and applications is vital in both cases. Also, constant visibility to application and network behaviors is required. With the increase in workload number and variety, quickly and efficiently resolving customer issues becomes increasingly difficult. Aviz Networks, together with Red Hat, provides solutions for workload and application lifecycle management with Red Hat OpenShift and Red Hat Enterprise Linux. For management of the network layer, they provide solutions at the Day-0 to Day-2 operations of AI fabrics. Aviz Service Node (ASN) provides a range of advanced capabilities such as deep packet inspection, packet-level dif...

AI's Hidden Bottleneck: The Network

Image
Networking for data centers is adapting to AI. Since GPU clusters are scaling up in size, they require very synchronized communication that can handle huge bandwidth as well as ultra-low latency. An Aviz Networks podcast features Taylor Allison, NVIDIA Senior Product Marketing Manager, who talks about the trend of networking with the current AI wave. During this conversation Allison touches upon how communication is essential in the AI training process, where GPUs transfer gradients back and forth, making network a crucial part of total workload performance. NVIDIA's Spectrum-X Ethernet, along with their InfiniBand options, has the capabilities to maintain performance in high demanding environments and help manage network related to AI. The interview further delves into NVIDIA's new Air feature that helps in deploying and configuring services by leveraging a Digital Twin feature, this can ultimately limit risk in planning for day 0 operations. Day 0 operations are by far the be...

What Happens Before a Spectrum-X AI Fabric Goes Live?

Image
The AI network doesn't just begin when the first GPU is plugged into a switch. A day zero in NVIDIA Spectrum-X involves topology definition, IP assignment, BGP setup, QoS settings, etc. At scale, the human effort that goes into configuring all of these components can turn into a substantial engineering effort. Intent-based automation can alleviate these difficulties. Starting with basic requirements such as Scale Unit and number of GPUs, IP ranges to utilize, Aviz ONES Fabric Designer can generate both the design of the fabric and configuration for the desired environment. This takes care of things on both ends of the system, from the IPCLOS underlay (leaf and spine setup, BGP configuration, ASN assignment, IP address, interfaces and MTU setup) all the way up to server-side setup (RoCEv2 parameters with PFC, and ECN, BlueField-3 settings, etc.). But setting up is only one hurdle. Validation is the other, so using NVIDIA AIR enables validation of the design, before rolling it out in...

Building High-Performance AI Fabrics with NVIDIA Spectrum-X

Image
Your AI network design needs more than just GPUs. Did you know that the Day 0 design decisions, topology, IP addressing, BGP, QoS, server connectivity, and more can have a dramatic impact on the performance of your network Fabric in production? Designing and deploying high-performance AI networking at scale using, for example, NVIDIA Spectrum-X is challenging. You can simplify Day 0 design and deployment with Aviz ONES Fabric Designer. Instead of hand-configure complex IPCLOS Fabrics and dealing with hundreds or even thousands of parameters, the ONES Fabric Designer lets teams input data at a high level – specifying the Scalable Unit (SU) size, the desired number of GPUs, and the IP address pool. The system then generates a tested and validated network Fabric design including spine and leaf configuration, BGP peering, interface settings, and IP assignments. Furthermore, the solution also fully automates the setup and configuration of NVIDIA BlueField and ConnectX SuperNICs including a...

Why AI's Biggest Opportunity Is Moving Beyond Hardware

Until recently, companies sold innovative hardware. Now companies are focused more on integrated and proprietary software. AI is disrupting companies' old ways. Quickly upgradable chips and networks are no longer the most important hardware to own. Long term value comes from data, models, and workflows that utilize AI to run applications. The biggest challenge to AI businesses is: Fast-developing hardware generations. Hard-to-keep-up-with upgrades in infrastructure. Dependence on hardware for developing applications. Growing demand for deploying AI quicker at a larger scale. The answer is not decreasing hardware spending. The solution is spending to develop hardware that is more flexible and evolving. The adaptability of infrastructure allows organizations to upgrade hardware without the difficulty of rebuilding applications. Integrating upgrades in hardware allows organizations to keep their AI workflows running and evolving. Implementing operating standards is what the market is ...

Can Today's Networks Handle Tomorrow's AI?

The way data center networks operate is changing with the evolving nature of AI workloads. In a recent Aviz Networks podcast, Taylor Allison, Senior Product Marketing Manager at NVIDIA, spoke about changes that will happen in networking when it comes to large-scale AI training and inference. To do AI training, GPUs need to work in tandem, which requires a network that is able to communicate with low latency and stay synchronized. Once you increase the number of GPUs, even more, the existing networking technologies will start to fail. The networking solutions from NVIDIA that are optimized for AI workloads include Spectrum-X built for AI clusters over Ethernet and InfiniBand, another high-performance networking architecture for AI overclusters. The podcast also covered NVIDIA Air and the digital twin technology. Before deploying configurations and automation to the production environments, users can create virtual environments in which they can test configurations and automation. This g...

How Aviz and Endace Strengthen Network Forensics with Continuous Packet Evidence

Image
To investigate cyberattacks and suspicious activity - and to meet compliance regulations - security teams rely on evidence found within packets. However, traffic is traversing not only across their data centers, but also cloud platforms, campuses, and edges, making it challenging to capture every essential packet. Some typical problems that teams encounter are: * Overloaded packet capture systems unable to handle traffic spikes. * Irrelevant data obscuring important evidence. * Inconsistent historical records creating drag on investigations. * Higher infrastructure costs, without an increase in visibility. The Aviz Networks-Endace integrated solution makes this task simpler by guaranteeing that only optimized, high-fidelity traffic arrives at packet capture systems. Aviz Deep Network Observability (DNO) smart services efficiently aggregate, filter, dedup, and normalize traffic before directing it to Endace Always-On Packet Capture, which ingests, indexes, and permanently stores the tra...

Why AI Factory Networking Should Be Validated Before Deployment

Image
Assembling AI infrastructure is significantly different than just installing servers and switches. Modern AI factories are highly interconnected and need to integrate networking, GPUs, storage, orchestration, and suites of applications. If configuration or integration issues appear in the deployed system, there will likely be major interruptions in the operation and costs associated with delays. Many companies struggle with: * A shortage of real-world physical lab testing. * Long validation/deployment cycles. * Discovering integration bugs at an inconvenient (read, costly!) late stage. * The hidden post-deployment rework needed to fix. That's where Aviz ONES on NVIDIA DSX Air helps deliver a better AI infrastructure build process: digital twin simulation. The idea is simple: instead of building, and then testing, why not design, simulate, and validate your entire AI networking environment beforehand? Engineers can create a model of their production-scale network, simulate traffic f...

Why Aviz Service Node v2.4 Is the Missing Intelligence Layer for Modern Networks

Image
In today's era, enterprise and telecom networks span much wider than a data center. These networks' traffic traverse through the cloud, edge, hybrid infrastructure, and wireless networks, giving network administrators a challenging time in gaining end-to-end visibility. Lack of smart traffic analysis has put security and network teams' fragmented views and increased incident response times. Typical issues faced are: The inability of getting visibility across a distributed environment. Replication of traffic which impacts data processing costs. Cost-intensive hardware appliances tied to proprietary technology. Lack of data to feed the AI-based observability platforms. Aviz Service Node v2.4 processes raw packet traffic to create machine-readable, application-aware metadata which is consumable by present-day observability platforms. ASN, through deep packet inspection (DPI) and application classification, enhances packet-aware traffic data with context so that AI/ ML can eff...

Why AI Factory Design Needs Real-Time Validation Instead of Physical Labs

Image
Modern AI infrastructure no longer only means networking together a stack of GPUs with a bunch of switches. The modern AI factory now relies on tightly optimized networking in which a slight configuration miss will affect performance, price or deployment timeline. Traditionally, this means it requires setting up and testing the network configuration in a physical lab which is expensive, time consuming and not easily scalable. Why, you ask? Because hardware testing alone can be cost prohibitive. Design cycles are serial and iterative, meaning designs and validation are split into distinct phases that take weeks or days to cycle through. When you're designing for the future, you have configuration misses which aren’t caught until you’ve rolled out production. It’s time to disrupt the cycle with digital twins. What NVIDIA AIR and Aviz ONES give you NVIDIA AIR provides an all-digital, virtualized AI network simulation where you can simulate all the switches, links, automation and confi...

Can Healthcare Stop Ransomware Faster Without Better Network Visibility?

Image
Many hackers are able to breach various parts of the healthcare system because the SecOps team is unable to quickly understand how far an attacker can spread throughout multiple systems. Hospitals, clinics, telemedicine, cloud services, devices and facility systems create traffic on a common network; however, due to lack of central control over how traffic flows & devices in the Silverado system, many of these facilities have blind spots pertaining to the analysis of East to West traffic, devices that are not monitored (unmanaged) & mixed infrastructure (devices that do not have agent-based security monitoring). Some examples of this are: Almost all types of ransomware operated by transferring themselves around to other locations before they perform file&folder encryptions. A lot of the different device types and workloads out there within the Healthcare ecosystem will not be able to utilize endpoint agents. Packet level evidence will assist with identifying unusual flows o...

Can Healthcare Device Security Improve Without Starting at the Network Traffic Layer?

Image
In today's world of healthcare, thousands of connected devices rely on one another, such as infusion pumps, imaging equipment, badge readers, controllers of H/VAC, clinical stations and tele-health endpoints. A majority of these types of devices do not support security agents, can't be patched and were not designed with cybersecurity in mind; they were simply designed to provide safe reliable clinical care.  Here are some real-world observations related to the above statement: • There are substantial amounts of valuable traffic signals that are generated on connected medial and facility devices • Many of the devices cannot be secured by the use of traditional endpoint controls • Packet-level visibility can provide insight into device behaviour patterns, communications, and risk • Optimized delivery of traffic helps improve the accuracy of asset discovery and monitoring tools • Network evidence can support security compliance and incident investigation processes There is a comm...

Can AI Networks Be Operated Before They Reach Production?

AI factories have already introduced a completely new way of thinking regarding networking. Network used to serve the purpose of connecting devices. Nowadays, it influences the performance of workload execution, the level of isolation, usage of GPU, troubleshooting speeds, etc. It has become possible to validate the operational processes by creating a digital twin environment before any changes get implemented. By doing that, it would be possible to test AI fabrics, topology configuration options, create tenants, telemetry gathering, and much more within the environment. It is important not to forget about switching from automated deployment and management to intelligent operations. It will be possible to analyze alarms, find reasons for network congestion, detect drifts in configuration, examine telemetry information, generate health summaries, and so on. In other words, there won't be such a need to use multiple dashboards and CLI configurations. To sum up, an AI-friendly network...

Is Continuous Network Evidence the Missing Link in PCI-DSS 4.0 Compliance?

Image
PCI-DSS 4.0 raises the bar for financial services organisations by making compliance more evidence-driven. Institutions now need to show ongoing proof for encryption in transit, certificate validity, user activity, network monitoring, insecure protocols, malware signals, and risk analysis across complex environments. This is difficult because financial traffic no longer stays inside one controlled data center. It moves across hybrid cloud, APIs, branches, partners, payment systems, and internal workloads. Application logs and endpoint telemetry are useful, but they often depend on the health and honesty of the system generating them. During failures or attacks, that evidence can become incomplete. Packet-derived evidence helps close this visibility gap by observing traffic independently at the network layer. It can show which protocols are being used, which TLS sessions are active, what certificates are present, which DNS queries are occurring, which web transactions are visible, and ...

Can AI Networks Be Operated Before They Reach Production?

Image
AI factories have already introduced a completely new way of thinking regarding networking. Network used to serve the purpose of connecting devices. Nowadays, it influences the performance of workload execution, the level of isolation, usage of GPU, troubleshooting speeds, etc. It has become possible to validate the operational processes by creating a digital twin environment before any changes get implemented. By doing that, it would be possible to test AI fabrics, topology configuration options, create tenants, telemetry gathering, and much more within the environment. It is important not to forget about switching from automated deployment and management to intelligent operations. It will be possible to analyze alarms, find reasons for network congestion, detect drifts in configuration, examine telemetry information, generate health summaries, and so on. In other words, there won't be such a need to use multiple dashboards and CLI configurations. To sum up, an AI-friendly network...

Can Public Sector Networks Keep Up With AI Without Increasing Complexity?

Image
IT departments in the public sector are under tremendous pressure to support cloud computing, cyber security, hybrid work, and AI solutions while adhering to shoestring budgets. An event in the industry revealed that using open networking helps make government networks adaptable and affordable. The key lesson? Avoid vendor lock-ins. With open networking based on standards, it is much easier for an organization to introduce new technology to its network without getting entangled with any hardware or software requirements. A solution to these concerns would be the deployment of artificial intelligence in the operations process. Instead of adding yet more technology and more people, intelligent observability makes the best use of what you have by converting alerts to information, accelerating issue resolution, and enabling multi-vendor environments to work effectively. The take away from all of this for county IT leaders would be that network modernization does not necessarily need to cos...

Can Financial Services Teams Prove DORA Resilience Without Packet-Level Evidence?

Image
DORA has changed operational resilience from a periodic compliance exercise into a daily responsibility for financial services teams. Banks, insurers, payment firms, investment firms, and ICT providers now need to show that risk management, incident response, third-party oversight, and monitoring are supported by clear evidence. One of the biggest challenges is not policy creation, but proof. Logs, quarterly attestations, and self-reported telemetry often leave gaps, especially during incidents or across hybrid cloud, data center, branch, and containerized environments. Some practical observations: • Continuous ICT monitoring must cover every asset and flow • Incident reconstruction needs evidence that survives a compromise • Third-party and AI service usage must be visible in real time • Encryption posture needs constant validation across systems • Detection must happen fast enough to meet reporting timelines The key takeaway is simple: DORA does not demand a specific technology. It d...

Why Does Healthcare Ransomware Keep Spreading Before Anyone Sees It?

Image
Ransomware in healthcare often succeeds before encryption even begins. Attackers usually enter through everyday paths such as phishing, exposed remote access, vulnerable edge systems, or compromised third parties. After that, the real damage happens quietly as they move across workloads, clinical systems, cloud environments, and connected devices looking for sensitive data and high-value systems. The challenge is visibility. Many healthcare networks include systems that cannot support agents, legacy medical devices, branch locations, cloud applications, and hybrid infrastructure. This makes it difficult for security teams to see East-West movement, unusual internal access, outbound exfiltration, and command-and-control behavior in time. Packet-derived metadata helps close this gap by capturing DNS, TLS, HTTP, flow, and session details directly from the network. This evidence strengthens ransomware detection and response workflows. Security platforms can use enriched network metadata to...

Can OT Security Improve Without Cleaner Network Visibility?

Image
Industrial and critical infrastructure environments are becoming more connected as IT, OT, and IoT systems work together. This improves efficiency, but it also increases security risk. For OT security tools to detect threats accurately, they need complete, clean, and relevant network traffic. A major challenge is that industrial networks often generate fragmented, duplicated, and noisy traffic across plants, substations, cloud-connected systems, and internal workloads. This makes it harder for security teams to identify real anomalies, monitor assets, and respond quickly without adding operational disruption. Deep network observability helps solve this by collecting traffic from different sources, filtering unnecessary data, removing duplicates, enriching packets with context, and delivering high-quality telemetry to security analytics systems. This allows teams to gain better asset visibility, improve threat detection accuracy, and reduce the load on monitoring tools. The key takeaway...

Can AI Factories Be Built Faster by Testing Before Hardware Arrives?

Image
I recently explored a discussion on how AI factory deployment is moving from a hardware-first model to a simulation-first approach. What stood out was how practical the shift feels for infrastructure teams that need to move faster without increasing deployment risk. AI factories are no longer simple GPU clusters. They include compute, networking, storage, security, orchestration, observability, and operations working together as one system. When teams wait for hardware before testing begins, issues often appear late in the lifecycle. That can lead to delays, rework, and lower confidence before production rollout. Some practical observations: • AI infrastructure is becoming too complex for traditional deployment methods • Simulation helps teams validate designs before physical systems are available • Connectivity, configuration, security, upgrade workflows, and failure scenarios can be tested earlier • Natural planning and validation workflows reduce dependency on late-stage trouble...