Startups · AI

ASIC Architect, Principal (AI Inference)

Positron AI · Remote (United States) · Remote

← All jobs
About Positron AI

About Positron AI Positron AI is building next-generation AI inference accelerators designed from the ground up for low-latency, high-throughput large language model inference. Backed by NEA.

About the role

We're looking for an exceptional Principal ASIC Architect to help define the future of AI inference hardware. You will collaborate across architecture, software, machine learning, systems, and silicon implementation to design next-generation AI accelerators capable of serving rapidly evolving frontier models. This role requires exceptional technical breadth, balancing mathematical understanding, ML trends, systems architecture, and practical silicon implementation.

What they're looking for

  • 15+ years of experience in ASIC architecture
  • Deep expertise in AI accelerator design and high-performance SoCs
  • Hands-on experience building performance models for memory systems and interconnects
  • Strong hardware/software co-design ability
  • Excellent written and verbal communication skills
More about this role

We're looking for an exceptional Principal ASIC Architect to help define the future of AI inference hardware. You will collaborate across architecture, software, machine learning, systems, and silicon implementation to design next-generation AI accelerators capable of serving rapidly evolving frontier models. This role requires exceptional technical breadth, balancing mathematical understanding, ML trends, systems architecture, and practical silicon implementation. This may be the single most important technical hire after the Chief Architect.

Architectural Exploration

  • Own architectural exploration across the evolving landscape of transformer models, MoE, sparse models, and long-context inference.
  • Evaluate emerging techniques including speculative decoding, continuous batching, KV-cache evolution, sub-quadratic and linear attention, and state-space models.
  • Assess memory architectures, communication architectures, and interconnects for next-generation accelerators.

Performance Modeling

  • Develop analytical models and build simulation frameworks to evaluate architectural tradeoffs.
  • Estimate latency, throughput, bandwidth, memory and compute utilization, scaling efficiency,...

Read the full posting on Positron AI's site ↗

Engineering

Build your edge while you search

Free tools for founders and investors, plus VC Unfiltered, our take on startups, venture and the people who build them.