About Positron AI Positron AI is building next-generation AI inference accelerators designed from the ground up for low-latency, high-throughput large language model inference. Backed by NEA.
About the role
Positron AI is looking for a Technical Product Manager to own AI inference and software technical product planning end to end. In this role, you will be the person who translates where models and inference systems are heading into concrete, well-scoped requirements for our inference software stack, spanning model coverage, numerics, inference-engine features and modes, serving-stack capabilities, and our managed service.
What they're looking for
- 10+ years experience with a strong mix of these types of technical skills:
- Production inference serving at scale touching on multi-tenancy, latency SLAs (TTFT/TPOT), batching and scheduling, disaggregated serving, KV-cache management, and observability
- The open-source inference ecosystem - runtimes (vLLM/SGLang-class), kernels, model ingestion, and how models are released, quantized, and adopted in practice
- Performance analysis spanning models & systems: utilization reasoning, tokens per $ and per W arithmetic, benchmark design - able to build and defend the math personally
- Competitive landscaping and analysis of inference providers and serving stacks, at a model lab, inference API provider, or AI hardware company
- 10+ years with a strong mix of these type of personal & analytic skills:
More about this role
Positron AI is looking for a Technical Product Manager to own AI inference and software technical product planning end to end. In this role, you will be the person who translates where models and inference systems are heading into concrete, well-scoped requirements for our inference software stack, spanning model coverage, numerics, inference-engine features and modes, serving-stack capabilities, and our managed service.
This is a deeply technical planning role that sits at the intersection of engineering, go-to-market, and the broader inference ecosystem. You will track the model frontier as a discipline, convert that movement into engineering requests before it becomes a customer escalation, and serve as the connective tissue between our engineering organization, our GTM teams, and our ecosystem partners. You will also be expected to use agentic AI daily as a core part of how the planning function operates.
- Be the leader in Positron for all aspects of AI Inference & Software Technical Product Planning.
- Write requirements for our inference software stack spanning model coverage, numerics, inference-engine features/modes, serving-stack features/modes, and managed-service...
Browse similar: AI jobs · AI startup jobs · Startup jobs · Remote jobs