Power AI workloads with AMD Instinct™ GPUs. Scale models faster with high-performance, dedicated cloud compute.
About the role
We’re seeking a Principal Network Engineer (L7) focused on owning and evolving large-scale, RoCEv2 data center networks powering next generation AI and ML infrastructure.
What they're looking for
- Bachelor’s degree in Computer Science, Electrical Engineering, or a related technical field, or equivalent practical experience
- Deep experience designing and operating RDMA and RoCEv2 networks in large-scale production environments supporting AI or HPC workloads
- Expert-level knowledge of switching hardware and their NOS, such as Arista, Juniper, and custom solutions using SONiC, including high-speed Ethernet fabrics
- Proven hands-on experience with congestion management and performance tuning using PFC, ECN, and DCQCN
- Strong experience with high-speed optics and cabling including 400G, 800G, and AEC, AOC, DAC, and structured cabling at scale
- Strong automation mindset, with experience using Python, Ansible, Terraform, Git, and production observability tooling
More about this role
Our mission is simple: deliver seamless, secure, reliable, and resilient AI compute at scale. We've built a versatile cloud platform that eliminates infrastructure barriers, empowering builders to focus on innovation instead of fighting their stack. Because breakthrough AI should move at the speed of ideas, not infrastructure.
We’re seeking a Principal Network Engineer (L7) focused on owning and evolving large-scale, RoCEv2 data center networks powering next generation AI and ML infrastructure.
You’ll work closely with our network architect and infrastructure leadership to define how the network is designed, implemented, and operated at scale, keeping over 8,000 GPUs burring today and scaling to cluster sizes reaching over 100,000 GPUs. You will be responsible for the architectural decisions that determine performance, reliability, and operational sanity at scale.
You’ll remain hands-on with high-speed optics, switching, routing, and congestion management in production clusters, while also setting the standards, patterns, and tooling other engineers build and operate against.
As a Principal Engineer, you own end-to-end network architecture, make high-impact design decisions, and...
Browse similar: Startup jobs · Remote jobs