Startups · AI

Member of Technical Staff, LLM Evaluation Infra

Inception Labs · San Mateo, United States · On-site

← All jobs
About Inception Labs

We are leveraging diffusion technology to develop a new generation of LLMs. Our dLLMs are much faster and more efficient than traditional autoregressive LLMs. Backed by AI Grant and Amplify.

About the role

We seek experienced engineers and scientists to develop the evaluation infrastructure and systems that drive frontier LLM performance. You'll design the frameworks that tell us whether our models are improving and ensure they perform reliably at scale in production.

What they're looking for

  • Build scalable, automated evaluation pipelines that integrate into model training and deployment workflows
  • Design, develop, and maintain robust evaluation frameworks and benchmarks for measuring LLM performance across diverse tasks and domains
  • Conduct rigorous statistical analysis of model outputs to identify failure modes, biases, and performance gaps
  • Partner with product and customer-facing teams to translate real-world use cases into meaningful evaluation criteria
  • Define and implement quantitative metrics that capture model quality, safety, reliability, and regression detection
  • BS/MS/PhD in Computer Science, Machine Learning, Statistics, or a related field (or equivalent experience)
More about this role

Inception creates the world’s fastest, most efficient AI models. Our Mercury model is the world’s fastest reasoning LLM and first commercially available diffusion LLM, delivering 5x greater speed and efficiency than today’s LLMs, with best-in-class quality.

We are the AI researchers and engineers behind such breakthrough AI technologies as diffusion models, flash attention, and DPO.

The Role

We seek experienced engineers and scientists to develop the evaluation infrastructure and systems that drive frontier LLM performance. You'll design the frameworks that tell us whether our models are improving and ensure they perform reliably at scale in production.

Key Responsibilities

  • Build scalable, automated evaluation pipelines that integrate into model training and deployment workflows.
  • Design, develop, and maintain robust evaluation frameworks and benchmarks for measuring LLM performance across diverse tasks and domains.
  • Conduct rigorous statistical analysis of model outputs to identify failure modes, biases, and performance gaps.
  • Partner with product and customer-facing teams to translate real-world use cases into meaningful evaluation criteria.
  • Define and implement...

Read the full posting on Inception Labs's site ↗

Build your edge while you search

Free tools for founders and investors, plus VC Unfiltered, our take on startups, venture and the people who build them.