# Member of Technical Staff - Inference at Hyperbolic

- Company: Hyperbolic
- What the company does: Hyperbolic is the open-access AI cloud made for AI developers, providing fast, affordable access to compute, inference, and AI services. Backed by Polychain Capital.
- Company website: https://hyperbolic.xyz
- Type: Startups (AI role)
- Level: Senior
- Location: San Francisco, CA
- Work setup: On-site
- Posted: 2026-09-09
- Apply by: 2026-10-24
- Apply: https://jobs.ashbyhq.com/hyperbolic/38121cf9-aa24-44f6-a500-ebcd2d3ca3ef
- Page: https://www.1752.vc/careers/jobs/hyperbolic-member-of-technical-staff-inference/

## About the role

We're looking for an Inference Engineer to build inference capabilities on top of Forge, our unified control plane, so customers can consume model tokens without managing GPUs and our NeoCloud partners get a full-stack path to their own token-factory offering. You'll own how models get deployed and served across clusters distributed around the world, on heterogeneous hardware.

## What they're looking for

- Strong general inference background with a broad, high-level command of the stack rather than a narrow specialty — you can reason about the whole path from request to token
- Deep Kubernetes experience, including hands-on ability to operate clusters in production, not just deploy to them
- Solid grasp of the concepts that govern inference performance: TTFT, disaggregated inference, speculative decoding, and KV cache and its inner workings
- Familiarity with modern inference frameworks and serving engines, and the judgment to evaluate and select among them for a given workload
- Working knowledge of NVIDIA Dynamo and how it fits into a distributed serving architecture
- Experience setting up monitoring, gateways, and endpoints for production inference services

Tags: Engineering
