Startups · AI

Founding Engineer -- AI Inference Stack

Piris Labs · San Francisco, CA, US · On-site

← All jobs
About Piris Labs

The Networking Layer of AI. Backed by Y Combinator.

About the role

Architect the Stack: Iterate on our software layer that orchestrates inference across heterogenous cluster of compute resources. Kernel and Compiler Optimization: Write and optimize high-performance kernels in CUDA, Triton, or custom targets to squeeze every drop of performance from the system.

What they're looking for

  • Curiosity-driven , with a genuine passion for compute architectures and problem solving
  • Systems Obsessed: You have a deep understanding of computer architecture, memory hierarchies, and low-level systems programming in C++, Rust, or CUDA
  • AI Fluent: You understand the guts of transformer architectures and have experience with inference frameworks like vLLM, TensorRT, ONNX, and Kubernetes
  • A First-Principles Thinker: You are not afraid to throw out the standard way of doing things if it means achieving a 10x performance gain
More about this role

The Role: We are looking for elite systems hackers who want to own the software layer of a new architecture. You will be writing kernels and orchestration logic that outperform existing solutions.

Architect the Stack: Iterate on our software layer that orchestrates inference across heterogenous cluster of compute resources.

Kernel and Compiler Optimization: Write and optimize high-performance kernels in CUDA, Triton, or custom targets to squeeze every drop of performance from the system.

The Runtime: Build the low-latency inference server, think a more performant custom version of vLLM or TensorRT-LLM, that manages KV cache at scale without the overhead of traditional PCIe bottlenecks.

Voice and Agentic Optimization: Solve the unique challenges of instant-on Voice AI, focused on latency, and the high-context demands of coding agents, focused on memory management.

Curiosity-driven , with a genuine passion for compute architectures and problem solving

Systems Obsessed: You have a deep understanding of computer architecture, memory hierarchies, and low-level systems programming in C++, Rust, or CUDA.

AI Fluent: You understand the guts of transformer architectures and have experience...

Read the full posting on Piris Labs's site ↗

Engineering

Build your edge while you search

Free tools for founders and investors, plus VC Unfiltered, our take on startups, venture and the people who build them.