AI-Native Cloud Platform. Backed by Index.
About the role
We’re looking to hire someone to own inference research hands-on and find ways to lower cost per token and latency on our customer workloads.
What they're looking for
- Systems or research background in LLM inference
- Deep understanding of LLM serving, from the kernel to the scheduler
- History of shipping products or research that people use in production-like scenarios, whether academic or industry
- Excited to collaborate closely with customers
- Enthusiasm for developer tools, cloud native technologies, and open source software
More about this role
Beam is an ultrafast AI inference platform. We built a serverless runtime that launches GPU-backed containers in less than 1 second and quickly scales out to thousands of GPUs. Developers use our platform to serve apps to millions of users around the globe. We're backed by Y Combinator, Tiger Global, and prominent developer-tool founders, including the founder of Snyk and former CTO of GitHub.
We’re looking to hire someone to own inference research hands-on and find ways to lower cost per token and latency on our customer workloads.
- Low-level inference optimization, from speculative decoding, quantization, KV-cache and memory management
- Work directly with customers to optimize their production workloads, and apply your learnings to our platform as product improvements
- High-level of autonomy to find the highest upside bets and guide the future of our inference platform based on your work
- Systems or research background in LLM inference
- Deep understanding of LLM serving, from the kernel to the scheduler
- History of shipping products or research that people use in production-like scenarios, whether academic or industry
- Excited to collaborate closely with customers
-...
Browse similar: AI jobs · AI startup jobs · Startup jobs · Remote jobs · San Francisco Bay Area