Agentic automation tailored to system-critical industries the world relies on. Book a free consultation today. Backed by a16z and a16z speedrun.
About the role
We're looking for a Lead / Senior / Staff QA Engineer to own quality for Nexxa's AI agent systems — products that plan, call tools, and take multi-step actions autonomously in industrial environments. This isn't traditional UI testing: you'll be designing evaluation frameworks for non-deterministic, tool-using systems, building golden datasets, catching regressions in reasoning quality, and stress-testing agent behavior under adversarial and real-world edge-case conditions.
What they're looking for
- 5+ years in QA/SDET roles, with demonstrated ownership of test strategy for complex systems
- Hands-on experience testing LLM-based products, chatbots, or AI agents — you understand why traditional deterministic test assertions break down for generative systems
- Practical experience with eval frameworks or tooling (e.g., promptfoo, DeepEval, RAGAS, LangSmith) or a track record of building your own
- Strong scripting/programming ability (Python preferred) to build test automation, data pipelines, and eval tooling
- Understanding of how LLM agents work: prompting, tool/function calling, context management, RAG, memory, and orchestration frameworks
- Experience designing test data and labeled datasets, including sourcing, sampling, and managing dataset drift over time
More about this role
Nexxa is building the best AI systems for heavy industries — enabling machines, systems and operations to think, decide and act autonomously across manufacturing, large-scale infrastructure, logistics and legacy environments.
Our mission is to translate deep technical breakthroughs into operational reality, solving some of the hardest systems-level problems in industry.
We're looking for a Lead / Senior / Staff QA Engineer to own quality for Nexxa's AI agent systems — products that plan, call tools, and take multi-step actions autonomously in industrial environments. This isn't traditional UI testing: you'll be designing evaluation frameworks for non-deterministic, tool-using systems, building golden datasets, catching regressions in reasoning quality, and stress-testing agent behavior under adversarial and real-world edge-case conditions.
You'll work closely with ML engineers, backend engineers, and Forward Deployed Engineers to define what "good" looks like for an agent operating in high-stakes industrial settings, then build the infrastructure and processes to measure it continuously.
Design and build evaluation harnesses and regression suites for LLM-based agents, covering...
Browse similar: AI jobs · AI startup jobs · Startup jobs · Remote jobs