Build what's next on the AI Native Cloud. Full-stack AI platform for inference, fine-tuning, and GPU clusters — powered by cutting-edge research. Backed by General Catalyst, Kleiner Perkins and NEA.
About the role
Together AI is building the Inference Platform that brings the most advanced generative AI models to the world. Our platform powers multi-tenant serverless workloads and dedicated endpoints, enabling developers, enterprises, and researchers to harness the latest LLMs, multimodal models, image, audio, video, and speech models at scale.
What they're looking for
- 5+ years of demonstrated experience building large-scale, fault-tolerant, distributed systems and API microservices
- Strong background in designing, analyzing, and improving efficiency, scalability, and stability of complex systems
- Excellent understanding of low-level OS concepts: multi-threading, memory management, networking, and storage performance
- Expert-level programming in one or more of: Rust, Go, Python, or TypeScript
- Knowledge of modern LLMs and generative models and how they are served in production is a plus
- Experience working with the open source ecosystem around inference is highly valuable, familiarity with SGLang, vLLM, or NVIDIA Dynamo will be especially handy
More about this role
Together AI is building the Inference Platform that brings the most advanced generative AI models to the world. Our platform powers multi-tenant serverless workloads and dedicated endpoints, enabling developers, enterprises, and researchers to harness the latest LLMs, multimodal models, image, audio, video, and speech models at scale.
If you get a thrill from optimizing latency down to the last millisecond, this is your playground. You’ll work hands-on with tens of thousands of GPUs (H100s, H200s, GB200s, and beyond), figuring out how to fully utilize every FLOP and every gigabyte of memory.
You’ll collaborate directly with research teams to bring frontier models into production, making breakthroughs usable in the real world. Our team also works closely with the open source community, contributing to and leveraging projects like SGLang, vLLM, and NVIDIA Dynamo to push the boundaries of inference performance and efficiency.
- Shape the core inference backbone that powers Together AI’s frontier models.
- Solve performance-critical challenges in global request routing, load balancing, and large-scale resource allocation.
- Work with state-of-the-art accelerators (H100s, H200s,...
Browse similar: AI jobs · AI startup jobs · Startup jobs · San Francisco Bay Area