Startups · AI

Software Engineer (C++ Systems)

Thunder Compute · San Francisco · On-site

← All jobs
About Thunder Compute

Thunder Compute makes GPUs abundant with GPU virtualization research, cloud infrastructure, and enterprise partnerships that unlock more data center capacity. Backed by Y Combinator.

About the role

Your work will focus on building the core C++ systems behind our virtualization layer. This includes low-latency performance optimization, distributed systems debugging, production reliability, and research into new techniques for improving GPU utilization. You will take ownership of complex systems from early experimentation through production deployment. Example projects may include:

What they're looking for

  • Exceptional modern C++ ability, including memory management, concurrency, performance optimization, and systems-level abstraction design
  • Deep understanding of operating systems, low-level networking, compilers, distributed systems, or computer architecture
  • Experience building and operating performance-critical C++ systems in production
  • Strong Linux systems programming and debugging ability
  • Ability to reason through unfamiliar systems across multiple layers of the stack
  • Strong work ethic and the ability to independently push a project from an experimental prototype through 100% completion under tight deadlines
More about this role

The world is building massive amounts of GPU capacity. Meanwhile, deployed GPUs are only 20% utilized.

This is because GPUs are not virtualized, while every other type of hardware is. For example CPUs and storage are allocated through virtual abstractions which efficiently manage the physical hardware, while GPUs are statically allocated on a one-to-one basis.

Thunder Compute is building this virtualization layer for GPUs. We have raised over $17.5M from Matrix Partners, Y Combinator, and leading angels from Coreweave, Microsoft, Cognition, and Anthropic.

Leading solutions for underutilization sit at the workload layer and are therefore only able to optimize specific use cases. We believe the ideal cluster optimization solution must be invisible to developers and compatible with all workloads; hence, it must sit at the systems layer.

We are a team of systems researchers productionizing cutting-edge GPU virtualization research to build this general-purpose optimization layer.

Concretely, our virtualization library abstracts GPUs across TCP networking. We use a userspace shim library, loaded through LD_PRELOAD , to intercept CUDA calls and send them over gRPC to a host server...

Read the full posting on Thunder Compute's site ↗

Engineering

Build your edge while you search

Free tools for founders and investors, plus VC Unfiltered, our take on startups, venture and the people who build them.