Startups · AI

Technical Lead, On-Device AI Inference

Hark · San Jose · On-site

← All jobs
About Hark

At Hark, we are building the most advanced personal intelligence in the world.

About the role

You'll own how Hark's models run on the silicon we ship: selecting the accelerators our devices are built around, co-designing architectures against real latency, memory, and power budgets, and building the low-level inference stack that turns a trained model into something that responds in milliseconds on a battery. You'll build and lead the team that does it. The ceiling on what our hardware can do is set here.

What they're looking for

  • 8–12+ years in high-performance computing, including production workloads deployed on GPUs, NPUs, or specialized accelerators
  • Deep understanding of attention, KV-cache behavior, quantization effects, and memory bandwidth limits
  • You've designed or optimized inference engines, distributed runtimes, or ML compilers, and you write the kernels yourself when it matters
  • Experience leading teams on performance-critical software. You've set direction on a stack, not just contributed to one
  • You've taken a model from a research checkpoint to running on constrained hardware in a product people use
More about this role

Hark is an artificial intelligence company building advanced, personalized intelligence. One that is proactive, multimodal, and capable of interacting with the world through speech, text, vision, and persistent memory.

We're pairing that intelligence with next-generation hardware to create a universal interface between humans and machines. While today's AI largely operates through chat boxes and decade-old devices, Hark is focused on what comes next: agentic systems that interact naturally with people and the real world.

To get there, we're developing multimodal models and next-generation AI hardware together - designed from the ground up as a single, unified interface for a new era of intelligent systems.

You'll own how Hark's models run on the silicon we ship: selecting the accelerators our devices are built around, co-designing architectures against real latency, memory, and power budgets, and building the low-level inference stack that turns a trained model into something that responds in milliseconds on a battery. You'll build and lead the team that does it. The ceiling on what our hardware can do is set here.

  • Evaluate GPUs, NPUs, DSPs, and specialized accelerators for...

Read the full posting on Hark's site ↗

On-Device Models

Build your edge while you search

Free tools for founders and investors, plus VC Unfiltered, our take on startups, venture and the people who build them.