Backed by NEA.
About the role
We are seeking an experienced Senior Software Engineer to contribute to the architecture, development, and scaling of a high-performance network monitoring and observability platform. This role will focus on building systems that provide deep visibility into RDMA, RoCE, InfiniBand, and TCP/IP networks. The ideal candidate has strong experience in distributed systems, Linux networking, and modern observability stacks (e.g., Grafana/Prometheus).
What they're looking for
- Bachelor's or Master's degree in Computer Science, Computer Engineering, Electrical Engineering, or a related technical field
- Strong hands-on programming experience in C++, Go, Python, Rust, or similar systems programming languages
- Proven experience delivering complex infrastructure projects and contributing to the design and implementation of distributed systems
- Experience building distributed systems, backend services, telemetry pipelines, or observability platforms
- Hands-on experience with RDMA, RoCE, InfiniBand, or other high-performance network fabrics
- Familiarity with libibverbs, RDMA verbs, RDMA CM, queue pairs, completion queues, memory registration, and related RDMA concepts
More about this role
Clockwork Systems was founded by Stanford researchers and veteran systems engineers who share a vision for redefining the foundations of distributed computing. As AI workloads grow increasingly complex, traditional infrastructure struggles to meet the demands of performance, reliability, and precise coordination. Clockwork is pioneering a software-driven approach to AI fabrics by delivering cross-stack observability to catch and quickly resolve problems, workload fault tolerance to keep jobs running through failures, and performance acceleration that dynamically routes and paces traffic to avoid congestion.
We are seeking an experienced Senior Software Engineer to contribute to the architecture, development, and scaling of a high-performance network monitoring and observability platform. This role will focus on building systems that provide deep visibility into RDMA, RoCE, InfiniBand, and TCP/IP networks. The ideal candidate has strong experience in distributed systems, Linux networking, and modern observability stacks (e.g., Grafana/Prometheus).
- Design, develop, and scale high-performance network monitoring platforms for RDMA, RoCE, InfiniBand, and TCP/IP infrastructure.
-...
Browse similar: Startup jobs · San Francisco Bay Area