Startups · AI

Research Engineer - Language Model Pre-Training

Zyphra · San Francisco · On-site

← All jobs
About Zyphra

The future of intelligence is open.

About the role

As a Research Engineer - Language Model Pre-Training , you'll shape our language model roadmap through end-to-end pretraining development. You will work extremely closely with our pretraining team, who will integrate your insights into our next-generation models.

What they're looking for

  • Strong engineering aptitude for rapidly implementing reliable and robust systems
  • Can rapidly learn new fields and are excited to implement new ideas
  • Excellent communication and collaboration skills, and can work effectively on both research and engineering implementation at scale
  • Deep expertise and intuition for solving machine learning problems and training models
  • Experience with training on large-scale (multi-node) GPU clusters
  • Deep understanding of model training pipelines – including model/data parallelism, distributed optimizers, etc
More about this role

As a Research Engineer - Language Model Pre-Training , you'll shape our language model roadmap through end-to-end pretraining development. You will work extremely closely with our pretraining team, who will integrate your insights into our next-generation models.

Large-scale training runs and model parallelization

Performance optimization of our pretraining stack

Dataset collection, processing, and evaluation

Architecture and methodology research, including optimizer ablations

Strong engineering aptitude for rapidly implementing reliable and robust systems

Can rapidly learn new fields and are excited to implement new ideas

Excellent communication and collaboration skills, and can work effectively on both research and engineering implementation at scale

Deep expertise and intuition for solving machine learning problems and training models

Experience with training on large-scale (multi-node) GPU clusters

Deep understanding of model training pipelines – including model/data parallelism, distributed optimizers, etc.

Strong grasp of proper experimental methodology for running rigorous ablations and other hypothesis testing

Understanding of large-scale, highly parallel data processing...

Read the full posting on Zyphra's site ↗

R&D - Engineering

Build your edge while you search

Free tools for founders and investors, plus VC Unfiltered, our take on startups, venture and the people who build them.