Startups · AI

Senior Machine Learning Engineer

TensorWave · Las Vegas, Nevada · On-site

← All jobs
About TensorWave

Power AI workloads with AMD Instinct™ GPUs. Scale models faster with high-performance, dedicated cloud compute.

About the role

We’re looking for a Senior Machine Learning Engineer to join our team during an exciting phase of growth. In this role, you’ll be responsible for building and operating the core systems that power large-scale ML training and inference across TensorWave’s GPU platform , working closely with cross-functional partners to support business objectives while upholding our standards for excellence, collaboration, and impact.

What they're looking for

  • Bachelor of Science in Computer Science, Computer Engineering, or a related technical field, or equivalent practical experience
  • Expertise supporting production ML systems using SLURM and Kubernetes
  • Strong understanding of GPU-accelerated workloads and distributed systems concepts
  • Solid Linux fundamentals and experience debugging infrastructure-level issues
  • Ability to build automation and tooling - Python, Go, etc
  • Experience working across schedulers, orchestration platforms, or cluster managers
More about this role

Our mission is simple: deliver seamless, secure, reliable, and resilient AI compute at scale. We've built a versatile cloud platform that eliminates infrastructure barriers, empowering builders to focus on innovation instead of fighting their stack. Because breakthrough AI should move at the speed of ideas, not infrastructure.

We’re looking for a Senior Machine Learning Engineer to join our team during an exciting phase of growth. In this role, you’ll be responsible for building and operating the core systems that power large-scale ML training and inference across TensorWave’s GPU platform , working closely with cross-functional partners to support business objectives while upholding our standards for excellence, collaboration, and impact.

Design, operate, and improve ML infrastructure systems supporting distributed training and inference workloads

Build reliable, repeatable workload execution and orchestration patterns across shared GPU environments

Troubleshoot performance, reliability, and scalability issues across the ML stack

Partner with ML, systems, and platform teams to improve developer experience and operational efficiency

Bachelor of Science in Computer Science,...

Read the full posting on TensorWave's site ↗

Engineering

Build your edge while you search

Free tools for founders and investors, plus VC Unfiltered, our take on startups, venture and the people who build them.