DriveNets, AMD Unveil AI Cluster Reference Architecture
End-to-end blueprint allows cost-efficient high-performance AI infrastructure based mostly on AMD Instinct
MI350 sequence GPUs and DriveNets AI Fabric
DriveNets, a pacesetter in large-scale networking options, at present introduced the publication of an end-to-end reference structure, that defines a validated design blueprint for constructing high-performance AI infrastructure utilizing AMD Instinct
MI350 sequence GPUs and DriveNets AI Fabric. The mixed resolution helps scale-out and scale-across architectures, together with front-end and storage networking based mostly on a single networking resolution. When mixed with automated orchestration and full-stack integration companies, it additionally helps speedy deployment and environment friendly end-to-end scaling.
The reference structure and the businesses’ rising strategic relationship replicate a shared imaginative and prescient for open, multi-vendor AI infrastructure. AMD just lately participated as a strategic investor in DriveNets’ $410 million Series D financing spherical, reinforcing a robust dedication to their joint prospects.
The reference structure, accessible right here, is accompanied by a complementary deployment information that outlines a system-level strategy for designing, deploying, and fine-tuning giant AI GPU clusters. Together, the paperwork cowl compute and networking design, in addition to end-to-end optimization throughout the full-stack, together with AMD ROCm
software program ecosystem, RCCL collective communications, community plugins, NICs, servers, and system-level orchestration to maximise coaching and inference efficiency.
Validated performances at scale
The reference structure consists of complete benchmarking throughout inference, coaching, community isolation, and cloth resiliency workloads on AMD Instinct MI355X GPU clusters – offering validated, reproducible proof of the platform’s readiness for production-scale AI deployments.
The structure consists of a number of community topology choices, together with detailed technical steering and finest practices for optimizing efficiency, scalability, isolation, and resiliency.
Benchmark outcomes present that DriveNets’ AI cloth delivers roughly 5% greater throughput and 10–15% decrease time to first token (TTFT) in comparison with publicly accessible business outcomes. At multi-node scale, the platform meets strict manufacturing service-level goals, together with sub-20 millisecond inter-token latency and a minimal of fifty output tokens per second per consumer.
Resiliency testing demonstrated steady collective communication efficiency underneath concurrent RDMA site visitors, with no degradation throughout transient hyperlink disruptions. Fabric bandwidth remained constant, with no observable restoration delays. The outcomes additionally display RCCL collective efficiency on AMD platforms similar to main publicly accessible NCCL benchmarks on equal configurations.
High efficiency with improved effectivity
These outcomes validate the joint structure as a scalable basis for the total LLM lifecycle, enabling excessive utilization and constant efficiency throughout coaching and inference workloads. The resolution is designed to maximise GPU effectivity, cut back whole price of possession, and ship sturdy cost-per-token efficiency, providing an open different to vertically built-in AI infrastructure stacks.
“AI infrastructure is shifting from single-vendor stacks to open, multi-vendor programs, and the community makes that shift work,” stated Ido Susan, co-founder and CEO of DriveNets. “This reference structure with AMD turns that right into a validated blueprint: AMD Instinct GPUs and our AI cloth ship greater throughput and decrease time to first token than publicly accessible outcomes — with the resiliency and consistency that manufacturing AI clusters require. This provides prospects constructing on AMD Instinct a validated, end-to-end blueprint to maximise GPU utilization and decrease cost-per-token, on open Ethernet.”
“Production AI clusters are outlined by how effectively compute, community and software program work collectively at scale,” stated Arvind Balakumar, company vice chairman, Global Cluster Engineering, AMD. “Our collaboration with DriveNets provides prospects a validated, open Ethernet reference structure that pairs AMD Instinct MI350 sequence GPUs with DriveNets AI Fabric to assist maximize GPU utilization, enhance workload effectivity and simplify deployment throughout coaching and inference workloads.”
A greater AI infrastructure alternative
The deployment information gives step-by-step directions overlaying compute, community interface card (NIC), networking, collective-communications libraries and workload setup, configurations, and efficiency optimizations.
AMD and DriveNets are intently collaborating to make sure end-to-end software program stack optimization, from NIC conduct and collective communications to system-level tuning. The corporations are working with joint prospects throughout manufacturing and proof-of-concept deployments, together with LLM builders and NeoCloud suppliers. The corporations additionally established a joint lab to allow prospects to run proof-of-concept (PoC) exams utilizing their very own workloads.
The submit DriveNets, AMD Unveil AI Cluster Reference Architecture first appeared on AI-Tech Park.
