Startups · AI

Senior Staff Software Engineer - Enterprise AI Platform

Mellanox · US, CA, Santa Clara · On-site

← All jobs
About Mellanox

NVIDIA end-to-end network management solutions enable monitoring, management, analytics and visibility. Backed by Sequoia.

About the role

We're building the platform that lets long-running autonomous agents operate safely inside NVIDIA's enterprise. These are not assistants on a developer's laptop. They are fleets of agents deployed in the cloud, running continuously at scale on shared accelerated compute. They take on real work across enterprise systems, so people get far more done than they could before.

What they're looking for

  • Design agent blueprints with clean interfaces for authorization, sandbox, memory, observability, and skills, so a new agent inherits its enterprise integrations from the platform
  • Enable composing and orchestrating agents: skills as first-class units with declarative manifests, multi-agent orchestration for delegation and handoff, and support for headless, long-running, autonomous agents
  • Broker credentials so multi-agent systems can authenticate and authorize without ever touching secrets, with least-privilege scoping on every token
  • Provide checkpoint and recovery so an agent resumes cleanly after a crash or restart
  • BS or MS in Computer Science, Engineering, or related field (or equivalent experience)
  • 12+ years building distributed systems, infrastructure, or developer platforms at scale
More about this role

We're building the platform that lets long-running autonomous agents operate safely inside NVIDIA's enterprise. These are not assistants on a developer's laptop. They are fleets of agents deployed in the cloud, running continuously at scale on shared accelerated compute. They take on real work across enterprise systems, so people get far more done than they could before. This role defines the constructs that agents are built from: the blueprints they start from, the tools, skills, and plugins that power them against enterprise data, the runtime safety harness that keeps them in bounds, and the connections into credential management, sandbox, memory, and observability. The team designs and ships these building blocks so that agent developers across the company can stand up a new agent, wire it in, and run it for days or weeks. Security and safe execution come out of the box, not something each team has to get right on its own.

Today an agent runs inside a single harness. Claude, Codex, and open-source agent harnesses each work differently underneath, with their own execution model, tool interface, and telemetry shape. The platform smooths over those differences, so a single skill,...

Read the full posting on Mellanox's site ↗

Build your edge while you search

Free tools for founders and investors, plus VC Unfiltered, our take on startups, venture and the people who build them.