Fast Models Optimized for Coding Agents. Backed by Y Combinator.
About the role
Morph was a 1 person company from 0 → 10M of revenue. We will be the first 10 person $10b company. You will work directly with the founders on problems that determine how efficiently frontier-scale models can be served. Small team, enormous compute, immediate production impact.
What they're looking for
- Have optimized complex production systems
- Can juggle 8+ Codex/Claude/other coding agents concurrently
- Understand GPU performance, memory bandwidth, collectives, and inference serving
- Are strong in Python, CuTEdsl, and comfortable navigating unfamiliar codebases
- Care about tokens per second, tokens per dollar, and correctness equally
More about this role
Morph was a 1 person company from 0 → 10M of revenue. We will be the first 10 person $10b company.
Every employee should contribute >30M of revenue/yr to the company.
The best candidates would be top 1% at multiple parts of the inference stack, yet have breadth across the whole stack.
Morph builds high margin inference infrastructure. Our stack spans kernels, model serving, routing, autoscaling, and capacity.
- Find the gap between theoretical hardware performance and production performance
- Trace latency and throughput regressions from the API layer down to individual kernels
- Optimize batching, scheduling, routing, quantization, and distributed execution
- Work on new research directions around caching
- Work with NVLink and RoCE
- Validate that every optimization preserves model quality and correctness
- Have optimized complex production systems
- Can juggle 8+ Codex/Claude/other coding agents concurrently
- Understand GPU performance, memory bandwidth, collectives, and inference serving
- Are strong in Python, CuTEdsl, and comfortable navigating unfamiliar codebases
- Care about tokens per second, tokens per dollar, and correctness equally
You will work directly with the...
Browse similar: AI jobs · AI startup jobs · Startup jobs · San Francisco Bay Area