Startups

Pre-Training Research Engineer

Sciforium · San Francisco, CA · On-site

← All jobs
About Sciforium

Sciforium builds the next generation of AI models with unprecedented efficiency, privacy, and versatility. Backed by SignalFire.

About the role

As a Pre-training Research Engineer, you’ll focus on model implementation, pertaining and scaling, and improving the quality of our byte-native and multimodal foundation models. You’ll build and iterate quickly on research ideas, contribute production-grade training code and infrastructure, and help deliver high-quality base models that can serve real-world use cases at scale.

What they're looking for

  • 5+ years of experience in machine learning research or engineering, with a proven track record of developing and pre-training large language or multimodal foundation models
  • Software Engineering: Strong general software engineering skills, with the ability to write robust and performant training code
  • ML Foundations: Solid understanding of deep learning fundamentals and modern pre-training methods and literature
  • Research and Experimentation: Ability to quickly implement research ideas and evaluate them using clear baselines, ablations, metrics, and analysis
  • GPU and Distributed Training: Hands-on experience running training workloads in GPU-based environments, with familiarity with distributed training
  • Education: MS in Computer Science, Machine Learning, Artificial Intelligence, Mathematics, or a related field
More about this role

Sciforium is an AI infrastructure company developing next-generation multimodal AI models and a proprietary, high-efficiency serving platform. Backed by multi-million-dollar funding and direct sponsorship from AMD with hands-on support from AMD engineers the team is scaling rapidly to build the full stack powering frontier AI models and real-time applications.

As a Pre-training Research Engineer, you’ll focus on model implementation, pertaining and scaling, and improving the quality of our byte-native and multimodal foundation models. You’ll build and iterate quickly on research ideas, contribute production-grade training code and infrastructure, and help deliver high-quality base models that can serve real-world use cases at scale.

Train large byte-native and multimodal foundation models across massive, heterogeneous corpora.

Implement and evaluate new model architectures, training objectives, and optimization methods.

Develop stable pre-training recipes and run scaling experiments for novel architectures.

Conduct ablations and analyze training dynamics, model behavior, and base-model quality.

Work with data and distributed training engineers to improve training efficiency,...

Read the full posting on Sciforium's site ↗

Engineering

Build your edge while you search

Free tools for founders and investors, plus VC Unfiltered, our take on startups, venture and the people who build them.