Startups · AI

AI/ML Infrastructure Engineer

Zensors · San Francisco · On-site

← All jobs
About Zensors

AI that can watch your business to create better customer experiences. Backed by Y Combinator.

About the role

Optimizing Core ML Pipelines: Identifying key bottlenecks in our current video analytics pipeline and performing in-depth analysis to ensure the best possible performance on current server and edge compute architectures. Cross-Stack Collaboration: Collaborating closely with AI research and platform engineering teams to optimize core parallel algorithms and influence the design of our next-generation inference infrastructure.

What they're looking for

  • BS/MS or Ph.D. in Computer Science, Electrical Engineering, or a related discipline
  • Strong programming skills in C/C++ and Python
  • Experience with model optimization , quantization, and efficient deep learning techniques (e.g., knowledge distillation, pruning)
  • Deep understanding of GPU hardware performance , including execution models, thread hierarchy, memory/cache management, and the cost/performance trade-offs of video processing
  • Experience with profiling and benchmarking tools (e.g., Nsight Systems, Nsight Compute) to validate performance on complex architectures
  • Experience identifying and resolving compute and data flow bottlenecks , particularly in high-bandwidth video processing pipelines
More about this role

The AI Infrastructure team at Zensors builds the engine that powers our visual sensing platform. We provide the tools to automate the lifecycle of our AI workflow, including model development, evaluation, optimization, deployment, and monitoring across thousands of video streams.

As a Machine Learning Engineer in ML Runtime & Optimization , you will develop technologies to accelerate the training and inference of computer vision models that power smart spaces and cities.

Optimizing Core ML Pipelines: Identifying key bottlenecks in our current video analytics pipeline and performing in-depth analysis to ensure the best possible performance on current server and edge compute architectures.

Cross-Stack Collaboration: Collaborating closely with AI research and platform engineering teams to optimize core parallel algorithms and influence the design of our next-generation inference infrastructure.

Model Acceleration: Applying advanced model optimization techniques—such as quantization (Int8/FP16), pruning, and layer fusion—to our Vision Transformers (ViTs) and CNNs to maximize throughput and minimize latency.

Building Efficient Operators: Working across the entire ML framework/compiler...

Read the full posting on Zensors's site ↗

Technical team

Build your edge while you search

Free tools for founders and investors, plus VC Unfiltered, our take on startups, venture and the people who build them.