Startups · AI

Software Engineer, Data Infrastructure

Cohere · New York · Remote

← All jobs
About Cohere

Backed by AI Grant.

About the role

Design, build, and operate the distributed storage system that feeds model training and evaluation. Run this system multiple on Kubernetes clusters at petabyte scale.

What they're looking for

  • Strong storage fundamentals, including replication, consistency, caching, and data lifecycle management
  • Strong coding ability. We work in Python and Go, experience in either is enough, but you should be willing to pick up the other
  • Experience running stateful systems on Kubernetes, including Persistent Volumes, CSI drivers, and StatefulSets
  • Hands-on experience with cloud object storage such as S3 as well as POSIX-style filesystems
  • Experience with parallel or HPC filesystems such as Weka, VAST, or Lustre
  • Familiarity with the data-loading and checkpointing patterns used in large-scale model training
More about this role

Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world business problems.

We’re training and deploying frontier models for enterprises who are building AI systems. We believe that our work is instrumental to the widespread adoption of AI and we are looking for folks that want to be part of that.

We obsess over what we build. Each one of us is responsible for contributing to increasing the capabilities of our models and the value they drive for our customers. Cohere is a team of researchers, engineers, designers, and more, who are all passionate about their craft.

We are a global technology company headquartered in Toronto with key offices in London, New York City, San Francisco, Montreal, Paris, Berlin and Seoul. Join us!

The Data Infrastructure team at Cohere is responsible for the storage and data movement layer underlying every model training run. We're building the unified storage layer that feeds our training workloads. It needs to serve petabytes of training data and model checkpoints fast enough to keep thousands of GPUs busy across several training clusters. In...

Read the full posting on Cohere's site ↗

Engineering & Infra

Build your edge while you search

Free tools for founders and investors, plus VC Unfiltered, our take on startups, venture and the people who build them.