We make the best chips physically possible for the large model needs of frontier labs.
About the role
Diagnose and fix issues across the OS, network, and cloud stack Reason about routing, DNS, firewalls, VPCs, private connectivity, and trust boundaries
What they're looking for
- We care more about instincts and pattern recognition than a checklist of tools. The right person has seen enough systems like ours to know which questions to ask
- Deep Linux systems knowledge. You can debug from userspace down to syscalls and routing tables, and you've spent enough time with namespaces, mounts, and process semantics to recognize their failure modes on sight
- Conservative about new patterns. When introducing a new module or tool, reads a few siblings first to pick up conventions. Spots and questions inherited patterns that don't apply to the new use case
- Threat-modeling instincts for shared infrastructure. Reasons about who can talk to what, what gets cached and trusted by whom, and the blast radius when something goes wrong. Distinguishes load-bearing security choices from defense-in-depth
- Operational thinking. Reasons about apply ordering, coordination windows, and "what fails first if X is misconfigured"
- Surgical git workflow. Knows the rebase tooling well enough that rewriting a branch isn't scary. Splits unrelated work into separate PRs. Never resorts to --no-verify or destructive shortcuts to make a problem go away
More about this role
We're a small engineering team designing a custom chip. The work is compute-heavy and tooling-heavy: hermetic builds, large verification jobs, custom developer environments, a self-hosted CI fleet, and a steadily growing collection of internal services that engineers depend on every day. The infrastructure that supports all of this — CI/CD, compute, shared filesystems, networking, internal tooling — already exists and has a system owner. We're hiring a second infrastructure engineer to broaden our capacity and add depth in areas adjacent to what we already have.
We're looking for a strong generalist with a network and systems bent. Someone who's comfortable debugging a Linux kernel issue in the morning, untangling a cloud networking problem at lunch, and writing a new MCP server for an unfamiliar protocol in the afternoon.
Diagnose and fix issues across the OS, network, and cloud stack
Reason about routing, DNS, firewalls, VPCs, private connectivity, and trust boundaries
Track down "permission denied" that's actually a mount option, or "build is slow" that's actually a metadata-server timeout
Improve, harden, and extend the network and host configuration we already have
Write...
Browse similar: Startup jobs · Remote jobs · San Francisco Bay Area