Startups · AI

Software Engineer, Senior

Positron AI · Remote (United States) · Remote

← All jobs
About Positron AI

About Positron AI Positron AI specializes in developing custom hardware systems to accelerate AI inference. Backed by NEA.

About the role

While this role is currently posted at a specific level, we are a growth-oriented organization and are open to hiring at a more senior level for the right candidate. Please note that this job description serves as a focused but generalized overview of the role; specific responsibilities and impact expectations will be tailored to the experience and seniority of the final hire.

What they're looking for

  • Design and implement high-performance inference software for LLMs on custom hardware
  • Develop and optimize C++-based libraries that efficiently utilize SIMD instructions, threading, and memory hierarchy
  • Work closely with FPGA and systems engineers to ensure efficient data movement and computational offloading between x86 CPUs and FPGAs
  • Optimize model execution via low-level optimizations, including vectorization, cache efficiency, and hardware-aware scheduling
  • Contribute to performance profiling tools and methodologies to analyze execution bottlenecks at the instruction and data flow levels
  • Apply NUMA-aware memory management techniques to optimize memory access patterns for large-scale inference workloads
More about this role

Role Overview

Senior Software Engineer – Machine Learning Systems & High-Performance LLM Inference

We are seeking a Senior Software Engineer to contribute to the development of high-performance software that powers execution of open-source large language models (LLMs) on our custom appliance. This appliance leverages a combination of FPGAs and x86 CPUs to accelerate transformer-based models. The software stack is written primarily in modern C++ (C++17/20) and heavily relies on templates, SIMD optimizations, and efficient parallel computing techniques.

Key Responsibilities

  • Design and implement high-performance inference software for LLMs on custom hardware.
  • Develop and optimize C++-based libraries that efficiently utilize SIMD instructions, threading, and memory hierarchy.
  • Work closely with FPGA and systems engineers to ensure efficient data movement and computational offloading between x86 CPUs and FPGAs.
  • Optimize model execution via low-level optimizations, including vectorization, cache efficiency, and hardware-aware scheduling.
  • Contribute to performance profiling tools and methodologies to analyze execution bottlenecks at the instruction and data flow levels.
  • Apply...

Read the full posting on Positron AI's site ↗

Engineering

Build your edge while you search

Free tools for founders and investors, plus VC Unfiltered, our take on startups, venture and the people who build them.