Sanas is pioneering the future of human communication. Founded by a team of Stanford researchers and entrepreneurs with deep industry experience, Sanas has developed the world's first real-time speech AI platform capable of accent translation, noise... Backed by General Catalyst, Insight and GV.
About the role
Sanas is bringing real-time speech and language models on-premise — deployed at scale directly inside sovereign data centers, not served from behind a hosted cloud endpoint. It's one of the most demanding environments in the industry: strict latency budgets, massive concurrency, and infrastructure that needs to be private and reliable.
More about this role
Sanas is bringing real-time speech and language models on-premise — deployed at scale directly inside sovereign data centers, not served from behind a hosted cloud endpoint. It's one of the most demanding environments in the industry: strict latency budgets, massive concurrency, and infrastructure that needs to be private and reliable.
We're looking for a deeply hands-on, experienced engineer to help lead that build. This is someone who shapes core infrastructure and architecture decisions rather than just executing against a specification, and who naturally raises the level of the engineers working alongside them.
Inference Optimizations
● Implement custom kernels and low-level optimizations to push the absolute limits of GPU compute.
● Apply graph optimization, operator fusion, and hardware-specific code generation at the ML compiler level.
● Profile, analyze, and resolve deep system bottlenecks to radically improve latency, throughput, and memory efficiency.
● Drive model-level execution improvements, including mixed precision and advanced quantization strategies.
Serving Optimizations
● Write and optimize custom inference server backends to handle complex business logic,...
Browse similar: AI jobs · AI startup jobs · Startup jobs · San Francisco Bay Area