Neurophos develops photonic AI processing technology that focuses on hardware solutions for accelerating artificial intelligence inference by replacing traditional electronic compute elements with Optical Processing Units (OPUs). Backed by Purple Arch Ventures.
About the role
We are seeking a performance engineer to own the benchmarking numbers behind the T100 optical inference accelerator. Architecture and product decisions here are made on measured performance and energy, and this role produces those figures for the same workloads at every level of fidelity we use: roofline and limiter analysis, architecture performance models, in-house RTL simulation, and measured runs on competing GPUs and accelerators.
What they're looking for
- BS or MS in Computer Engineering, Electrical Engineering, Computer Science, or equivalent practical experience
- 5+ years of experience in GPU performance engineering, accelerator benchmarking, HPC performance measurement, or ML systems measurement
- Track record of building or operating benchmark harnesses that produced measured results on real GPUs or accelerators, including turning a Hugging Face model card, paper, or application description into a runnable benchmark
- Hands-on experience with roofline analysis, limiter analysis, or analytical performance modeling
- GPU performance analysis with NVIDIA Nsight Systems and Nsight Compute, or an equivalent profiler, covering HBM-bound versus compute-bound analysis, precision (FP16, BF16, FP8, INT8), and batching
- Working knowledge of LLM inference stacks such as Hugging Face, vLLM, SGLang, or TensorRT-LLM, including prefill versus decode, continuous batching, and MoE
More about this role
The demand for new data centers and AI compute is rapidly outpacing the planet's energy capacity. Digital solutions are hitting a power wall as we approach the physical limits of traditional silicon. Conquering this bottleneck means rethinking the fundamental architecture of inference compute. The industry's current path can't meet the need, so we're taking a different approach.
Instead of traditional electronic circuits, we use silicon photonics and an active, programmable metasurface to perform matrix multiplications at the speed of light. Our optical cells are 10,000x smaller than traditional photonic components, enabling unprecedented density. By using photonics instead of electricity, our chips become more efficient as they scale. This architecture will deliver up to 100 times the energy efficiency of existing solutions while significantly improving performance for large-scale AI inference.
We’ve assembled a world-class team of industry veterans and recently raised a $110M Series A led by Gates Frontier. Participants include M12 (Microsoft’s Venture Fund), Carbon Direct Capital, Aramco Ventures, Bosch Ventures, Tectonic Ventures, Space Capital, and others.
Join us and shape...
Browse similar: Startup jobs