Startups

Staff Software Engineer, Ray Data

Anyscale · San Francisco · On-site

← All jobs
About Anyscale

Powered by Ray, Anyscale helps AI builders run data-intensive workloads to build and deploy Foundation Models and AI at scale on any cloud. Backed by NEA, a16z and Amplify.

About the role

Design, build, and improve the core systems that power Ray Data , with a focus on performance, scalability, and reliability. Design and optimize distributed execution across different stages of data pipelines in heterogeneous environments.

What they're looking for

  • 6+ years of experience building production-grade software, infrastructure, or developer-facing systems, with strong Python engineering experience
  • 6+ years of experience personally owning core architectural decisions within a distributed data or compute engine, rather than primarily operating or using a platform someone else designed
  • Deep experience with distributed systems internals, such as scheduling, fault tolerance, data partitioning, distributed execution, performance optimization, or database and query engine internals
  • A track record of reasoning through system-level tradeoffs and defending architectural decisions, such as batch vs. streaming, static vs. dynamic resource allocation, or consistency vs. availability
  • Passion for solving the unsolved problems in large-scale AI infrastructure and building systems that enable the next generation of AI applications
More about this role

At Anyscale , we're on a mission to democratize distributed computing and make it accessible to software developers of all skill levels. We’re commercializing Ray , a popular open-source project that's creating an ecosystem of libraries for scalable machine learning. Companies like OpenAI , Uber , Spotify , Instacart , Cruise , and many more, have Ray in their tech stacks to accelerate the progress of AI applications out into the real world.

With Anyscale, we’re building the best place to run Ray, so that any developer or data scientist can scale an ML application from their laptop to the cluster without needing to be a distributed systems expert.

Proud to be backed by Andreessen Horowitz, NEA, and Addition with $250+ million raised to date.

Ray Data is a Python-native data processing engine and a one-stop shop for all AI data processing needs. Ray Data provides performant, first-class integration with cutting-edge AI frameworks using both multimodal and structured data.

The Ray Data team develops and maintains Ray Data , building the underlying distributed data processing infrastructure that powers modern AI workloads. We are a team of engineers passionate about solving...

Read the full posting on Anyscale's site ↗

Engineering

Build your edge while you search

Free tools for founders and investors, plus VC Unfiltered, our take on startups, venture and the people who build them.