Startups · AI

ML Engineer, Inference & Optimization

Pika · Palo Alto HQ · On-site

← All jobs
About Pika

AI creative tools. Built for Creatives. Backed by Lightspeed and AI Grant.

About the role

We are seeking Senior/Staff level Inference Engineers to accelerate the performance of Pika's AI-driven products. In this highly technical role, you will operate at the intersection of cutting-edge inference acceleration, GPU parallelism, advanced model deployment, and video generation technologies. Your expertise will drive significant improvements to model speed and efficiency, ensuring our creative AI systems deliver industry-leading user experiences at scale.

What they're looking for

  • Experience : 5+ years engineering experience, with a strong track record in inference acceleration and model deployment at scale
  • Inference Mastery : Proven expertise in inference optimization, including quantization, attention acceleration, and deep learning compiler stacks
  • GPU & Parallelism : Deep knowledge of GPU programming (CUDA, NCCL) and experience with SP, TP, PP, and other forms of parallelism for distributed inference
  • AI Domain Knowledge : Familiarity with video generation (videogen) models and large language models (LLMs)
  • Collaboration : Strong cross-discipline communication skills, able to drive shared goals across research and engineering functions
  • Ownership Mindset : Self-driven, solutions-oriented, and capable of managing ambiguity in a fast-paced startup environment
More about this role

We are seeking Senior/Staff level Inference Engineers to accelerate the performance of Pika's AI-driven products. In this highly technical role, you will operate at the intersection of cutting-edge inference acceleration, GPU parallelism, advanced model deployment, and video generation technologies. Your expertise will drive significant improvements to model speed and efficiency, ensuring our creative AI systems deliver industry-leading user experiences at scale.

You will design and optimize inference pipelines, implement state-of-the-art acceleration techniques, and work closely with researchers and engineers across the team to push the boundaries of what’s possible in real-time AI deployment. Your efforts will play a foundational role in powering the next generation of Pika’s video and language models.

Accelerate Inference : Lead and implement advanced inference acceleration techniques, including attention optimization and quantization for efficient model serving.

Maximize GPU Parallelism : Engineer and optimize GPU strategies across tensor, sequence, and pipeline parallelism (TP, SP, PP) for maximal efficiency and scalability.

Programming for Performance : Develop and optimize...

Read the full posting on Pika's site ↗

Research

Build your edge while you search

Free tools for founders and investors, plus VC Unfiltered, our take on startups, venture and the people who build them.