About Positron AI Positron AI specializes in developing custom hardware systems to accelerate AI inference. Backed by NEA.
About the role
While this role is currently posted at a specific level, we are a growth-oriented organization and are open to hiring at a more senior level for the right candidate. Please note that this job description serves as a focused but generalized overview of the role; specific responsibilities and impact expectations will be tailored to the experience and seniority of the final hire.
What they're looking for
- Design and implement high-performance inference software for LLMs on custom hardware
- Develop and optimize C++-based libraries that efficiently utilize SIMD instructions, threading, and memory hierarchy
- Work closely with FPGA and systems engineers to ensure efficient data movement and computational offloading between x86 CPUs and FPGAs
- Optimize model execution via low-level optimizations, including vectorization, cache efficiency, and hardware-aware scheduling
- Contribute to performance profiling tools and methodologies to analyze execution bottlenecks at the instruction and data flow levels
- Apply NUMA-aware memory management techniques to optimize memory access patterns for large-scale inference workloads
More about this role
Role Overview
Senior Software Engineer – Machine Learning Systems & High-Performance LLM Inference
We are seeking a Senior Software Engineer to contribute to the development of high-performance software that powers execution of open-source large language models (LLMs) on our custom appliance. This appliance leverages a combination of FPGAs and x86 CPUs to accelerate transformer-based models. The software stack is written primarily in modern C++ (C++17/20) and heavily relies on templates, SIMD optimizations, and efficient parallel computing techniques.
Key Responsibilities
- Design and implement high-performance inference software for LLMs on custom hardware.
- Develop and optimize C++-based libraries that efficiently utilize SIMD instructions, threading, and memory hierarchy.
- Work closely with FPGA and systems engineers to ensure efficient data movement and computational offloading between x86 CPUs and FPGAs.
- Optimize model execution via low-level optimizations, including vectorization, cache efficiency, and hardware-aware scheduling.
- Contribute to performance profiling tools and methodologies to analyze execution bottlenecks at the instruction and data flow levels.
- Apply...
Browse similar: AI jobs · AI startup jobs · Startup jobs · Remote jobs