Create AI videos from your ideas using HeyGen. Input text, image, or audio to create complete videos with narration, captions, visuals, and animations. Backed by Conviction.
About the role
HeyGen is building AI applications including Avatar IV, Photo Avatar, Interactive Avatar, and Video Translation. We’re looking for a Software Engineer focused on GPU performance to make the systems behind these experiences faster and more efficient. You will work across model execution and inference infrastructure, using profiling and measurement to improve latency, throughput, and GPU cost. This role is a fit for an engineer who enjoys understanding how software uses the hardware beneath it. Found on 1752vc Careers, the job board for startup and VC roles.
What they're looking for
- Experience optimizing GPU-based AI workloads or high-performance computing systems
- Proficiency in Python and experience with PyTorch or a similar machine learning framework
- Strong curiosity about GPU hardware, including memory bandwidth, cache behavior, tensor cores, and data movement between CPU and GPU
- Experience using profiling tools to connect hardware behavior to application-level bottlenecks and validate improvements
- Ability to turn performance experiments into reliable production changes and communicate tradeoffs clearly
More about this role
At HeyGen, our mission is to make visual storytelling accessible to all. Over the last decade, visual content has become the preferred method of information creation, consumption, and retention. But the ability to create such content, in particular videos, continues to be costly and challenging to scale. Our ambition is to build technology that equips more people with the power to reach, captivate, and inspire audiences.
Learn more at www.heygen.com . Visit our Mission and Culture doc here .
HeyGen is building AI applications including Avatar IV, Photo Avatar, Interactive Avatar, and Video Translation. We’re looking for a Software Engineer focused on GPU performance to make the systems behind these experiences faster and more efficient.
You will work across model execution and inference infrastructure, using profiling and measurement to improve latency, throughput, and GPU cost. This role is a fit for an engineer who enjoys understanding how software uses the hardware beneath it.
- Use NVIDIA Nsight Systems, Nsight Compute, and PyTorch Profiler to investigate GPU utilization, kernel execution, memory bandwidth, and CPU–GPU data movement.
- Identify bottlenecks across model...
Browse similar: AI jobs · AI startup jobs · Startup jobs · San Francisco Bay Area