Startups · AI

Staff Software Engineer, AI Inference Gateway

Salad · Remote (United States) · Remote

← All jobs
About Salad

Our mission is to democratize cloud computing by utilizing latent consumer resources to power a sustainable, affordable, environmentally friendly cloud for everyone. salad.com/about. Backed by Techstars.

About the role

SaladCloud runs inference on thousands of geodistributed workstation and consumer GPUs — NVIDIA and AMD — that nobody else can use. Sitting in front of that fleet is a unique AI Gateway written in Rust, built on Pingora and WireGuard, that serves real-time and batch inference and embeddings to customers as a single, reliable API. You'll own it.

What they're looking for

  • Strong production Rust experience, ideally on high-throughput networked services
  • Hands-on experience running LLM inference servers (vLLM, llama.cpp, TGI, TensorRT-LLM, or similar) — you know what a KV cache is and why it fills up
  • Distributed systems fundamentals: load balancing, backpressure, failure handling, unreliable nodes
  • Comfort operating what you build: metrics, tracing, on-call instincts
  • Willing to be on-call for a system you helped make quiet — we treat after-hours pages as a design bug to fix, not a lifestyle
  • Clear written and verbal communication, you can run a design without hand-holding and bring others along
More about this role

Salad Technologies · Remote · Full-time

SaladCloud runs inference on thousands of geodistributed workstation and consumer GPUs — NVIDIA and AMD — that nobody else can use. Sitting in front of that fleet is a unique AI Gateway written in Rust, built on Pingora and WireGuard, that serves real-time and batch inference and embeddings to customers as a single, reliable API. You'll own it.

This is an infrastructure role with real autonomy. You'll design and ship the systems that decide how requests are fulfilled, when to add or drop capacity, and how to bill for every token — and you'll be the person who knows whether a new model, quantization, or GPU is actually worth running.

  • Own the AI Gateway system end to end: request routing, streaming, batch/async job handling, and horizontal scaling across multiple gateway servers
  • Build and tune the fleet-efficiency algorithms — autoscaling nodes, scoring performance, and evicting underperformers so the network converges on peak throughput per dollar
  • Operate and scale the underlying inference nodes running vLLM, llama.cpp, and similar servers
  • Design and maintain OpenTelemetry-based observability across the gateway and the fleet;...

Read the full posting on Salad's site ↗

SaladCloud

Build your edge while you search

Free tools for founders and investors, plus VC Unfiltered, our take on startups, venture and the people who build them.