NUVACORE builds for the stratosphere. Maximum performance. Absolute area efficiency. No compromise. A general-purpose CPU core designed to excel everywhere, from core infrastructure to advanced AI systems, including the continuous demands of agentic computing. Backed by Sequoia.
About the role
Design, develop and own Nuvacore’s cycle-accurate CPU performance simulator — its architecture, fidelity, and long-term roadmap — as the authoritative reference model for the design team. Experience developing trace-based and execution-driven CPU performance simulators. Lead microarchitectural exploration across all major CPU domains: frontend (fetch, decode, branch prediction), out-of-order engine (rename, dispatch, execution), and memory subsystem (caches, prefetchers, TLBs, coherence).
What they're looking for
- MS in Computer Architecture, Computer Engineering, or related field (PhD preferred)
- 20+ years (Lead) or 8+ years (ICs) in CPU microarchitecture and/or performance engineering
- 20+ years (Lead) or 8+ years (ICs) of hands-on experience with cycle-accurate C++ CPU performance simulators (e.g. gem5 or equivalent in-house tools)
- Deep expertise in at least one of: branch prediction, fetch/decode, rename/dispatch, OOO execution, memory subsystem, or prefetchers
- Strong C++ and Python / Perl skills, ability to write and analyse assembly for microarchitectural test cases
- Proven track record driving performance from pathfinding through silicon (RTL correlation and/or post-Si debug)
More about this role
Nuvacore is building ground-up CPU silicon for next-generation compute workloads. As our CPU Performance Modeling Lead, you will own the performance modeling infrastructure and methodology that drives every architectural decision we make — from early pathfinding through tape-out and post-silicon correlation.
Design, develop and own Nuvacore’s cycle-accurate CPU performance simulator — its architecture, fidelity, and long-term roadmap — as the authoritative reference model for the design team. Experience developing trace-based and execution-driven CPU performance simulators.
Lead microarchitectural exploration across all major CPU domains: frontend (fetch, decode, branch prediction), out-of-order engine (rename, dispatch, execution), and memory subsystem (caches, prefetchers, TLBs, coherence).
Drive model-to-RTL correlation through all execution phases; own debug and resolution of performance miscorrelation between the performance model, RTL simulation, and post-silicon results.
Collaborate closely with design, verification, compiler, and system software teams to align architecture decisions with implementation constraints and software stack realities.
Mentor engineers across the...
Browse similar: AI jobs · AI startup jobs · Startup jobs