About Wherobots Wherobots is the AI Context Engine for the Physical World: the missing infrastructure layer for AI that needs to reason about our physical reality. Backed by Felicis.
About the role
This is a distributed-systems-first role with meaningful ML infrastructure ownership. You will spend most of your time building high-throughput, GPU-aware data pipelines that turn massive raster archives into features, predictions, and published outputs at global scale. The role sits at the intersection of distributed systems, ML inference, and geospatial data infrastructure.
What they're looking for
- Design and operate end-to-end ML pipelines : Build pipelines over massive raster archives such as Zarr and COG, from ingestion to feature generation to inference to publication
- Build high-throughput distributed pipelines : Use Ray (Datasets and actors) with careful control over I/O, compute overlap, and backpressure to keep clusters fully utilized
- Optimize GPU inference at scale : Tune PyTorch inference pipelines using batching, CUDA stream overlap, and memory-aware scheduling to maximize throughput per GPU
- Develop spatial data processing patterns : Implement tiling, overlapping windows, and accumulators that match the access patterns of spatial models
- Ensure production reliability : Build in retries, checkpointing, observability, and cost-efficient scaling so long-running global jobs are debuggable and resilient to failure
- Build reusable platform abstractions : Collaborate on abstractions that generalize across datasets, models, and product use cases so new workflows ship quickly
More about this role
About the role
Wherobots is looking for a passionate, skilled, and experienced Machine Learning Engineer to help architect, build, and operate the large-scale geospatial ML platform that powers GeoAI workflows on hundreds of terabytes to petabytes of raster data.
This is a distributed-systems-first role with meaningful ML infrastructure ownership. You will spend most of your time building high-throughput, GPU-aware data pipelines that turn massive raster archives into features, predictions, and published outputs at global scale. The role sits at the intersection of distributed systems, ML inference, and geospatial data infrastructure. If you can design clean dataflow, get the most out of a GPU cluster, and turn research prototypes into resilient production systems, we should talk.
We are 100% cloud-native and build our product using modern, reliable tooling. We use Ray, PyTorch, and the scientific Python stack (PyArrow, NumPy, Xarray) to operate on Zarr, Cloud-Optimized GeoTIFF (COG), GeoParquet, and Parquet data on object storage.
If you are passionate about building cutting-edge ML infrastructure for the physical world and want to be part of a fast-growing company at the...
Browse similar: AI jobs · AI startup jobs · Startup jobs