The Networking Layer of AI. Backed by Y Combinator.
About the role
Architect the Stack: Iterate on our software layer that orchestrates inference across heterogenous cluster of compute resources. Kernel and Compiler Optimization: Write and optimize high-performance kernels in CUDA, Triton, or custom targets to squeeze every drop of performance from the system.
What they're looking for
- Curiosity-driven , with a genuine passion for compute architectures and problem solving
- Systems Obsessed: You have a deep understanding of computer architecture, memory hierarchies, and low-level systems programming in C++, Rust, or CUDA
- AI Fluent: You understand the guts of transformer architectures and have experience with inference frameworks like vLLM, TensorRT, ONNX, and Kubernetes
- A First-Principles Thinker: You are not afraid to throw out the standard way of doing things if it means achieving a 10x performance gain
More about this role
The Role: We are looking for elite systems hackers who want to own the software layer of a new architecture. You will be writing kernels and orchestration logic that outperform existing solutions.
Architect the Stack: Iterate on our software layer that orchestrates inference across heterogenous cluster of compute resources.
Kernel and Compiler Optimization: Write and optimize high-performance kernels in CUDA, Triton, or custom targets to squeeze every drop of performance from the system.
The Runtime: Build the low-latency inference server, think a more performant custom version of vLLM or TensorRT-LLM, that manages KV cache at scale without the overhead of traditional PCIe bottlenecks.
Voice and Agentic Optimization: Solve the unique challenges of instant-on Voice AI, focused on latency, and the high-context demands of coding agents, focused on memory management.
Curiosity-driven , with a genuine passion for compute architectures and problem solving
Systems Obsessed: You have a deep understanding of computer architecture, memory hierarchies, and low-level systems programming in C++, Rust, or CUDA.
AI Fluent: You understand the guts of transformer architectures and have experience...
Browse similar: AI jobs · AI startup jobs · Startup jobs · Founding team roles · San Francisco Bay Area