Build what's next on the AI Native Cloud. Full-stack AI platform for inference, fine-tuning, and GPU clusters — powered by cutting-edge research. Backed by General Catalyst, Kleiner Perkins and NEA.
About the role
As a Forward Deployed Engineer (FDE) focused on Inference & Post-Training, you will be a hands-on technical partner to our most strategic customers — production AI teams looking to leverage high quality models and do inference at scale. For us, FDE is not a replacement for a Solutions Architect; you will partner with our SAs as a deep-domain specialist in inference optimization, fine-tuning pipelines, and production deployment.
What they're looking for
- Experience: 5+ years in a technical role, with a strong focus on inference systems, open-source LLM deployment, or post-training workflows
- Inference Engine Depth: Expert-level, hands-on experience with inference engines (e.g., vLLM, TensorRT-LLM, SGLang), ability to diagnose and resolve performance issues at the engine level
- Inference Optimization: Deep knowledge of KV cache tuning, speculative decoding, tensor parallelism, pipeline parallelism, and quantization techniques
- Post-Training Knowledge: Hands-on experience with fine-tuning and post-training pipelines, including LoRA, SFT, DPO, RLHF, and GRPO, ability to advise on system design
- Model Landscape Awareness: Broad knowledge of state-of-the-art open-source models and strong judgment on model selection for specific customer use cases, hardware profiles, and performance targets
- Coding Proficiency: Strong Python skills, comfortable working in production environments
More about this role
As a Forward Deployed Engineer (FDE) focused on Inference & Post-Training, you will be a hands-on technical partner to our most strategic customers — production AI teams looking to leverage high quality models and do inference at scale. For us, FDE is not a replacement for a Solutions Architect; you will partner with our SAs as a deep-domain specialist in inference optimization, fine-tuning pipelines, and production deployment. As key contributors to both the CX, Engineering, and Sales organizations, FDEs add tremendous value by ensuring we can meet the requirements of our most complex POCs, facilitate successful platform adoption, and guide tailored optimization efforts — directly impacting customer success, company growth, and the hardening of our core platform.
- Inference Engine Optimization: Select, configure, and optimize inference engine based on hardware, model architecture, and workload profile
- Configuration & Performance Tuning: Develop configuration updates to win critical POCs, benchmarks, and optimize customer deployments; tune KV cache, apply speculative decoding, determine optimal tensor parallelism, and determine quantization strategy to hit throughput and...
Browse similar: AI jobs · AI startup jobs · Startup jobs · San Francisco Bay Area