# Technical Lead, On-Device AI Inference at Hark

- Company: Hark
- What the company does: At Hark, we are building the most advanced personal intelligence in the world.
- Company website: https://hark.com
- Type: Startups (AI role)
- Level: Senior
- Location: San Jose
- Work setup: On-site
- Pay: $300K to $500K base salary per year (USD)
- Posted: 2026-09-01
- Apply by: 2026-10-16
- Apply: https://job-boards.greenhouse.io/hark/jobs/4392086009
- Page: https://www.1752.vc/careers/jobs/hark-technical-lead-on-device-ai-inference/

## About the role

You'll own how Hark's models run on the silicon we ship: selecting the accelerators our devices are built around, co-designing architectures against real latency, memory, and power budgets, and building the low-level inference stack that turns a trained model into something that responds in milliseconds on a battery. You'll build and lead the team that does it. The ceiling on what our hardware can do is set here.

## What they're looking for

- 8–12+ years in high-performance computing, including production workloads deployed on GPUs, NPUs, or specialized accelerators
- Deep understanding of attention, KV-cache behavior, quantization effects, and memory bandwidth limits
- You've designed or optimized inference engines, distributed runtimes, or ML compilers, and you write the kernels yourself when it matters
- Experience leading teams on performance-critical software. You've set direction on a stack, not just contributed to one
- You've taken a model from a research checkpoint to running on constrained hardware in a product people use

Tags: On-Device Models
