Startups

2027 Internship Evaluation Engineer, Metric Prototyping

Bedrock Robotics · San Francisco, CA · On-site

← All jobs

About the role

Every time Bedrock works with a new machine for tasks like grading pads, digging trenches, or loading trucks, we need a way to verify the machine is actually performing well. The Evaluation team defines what "good" looks like: building the metrics, benchmarks, and evaluation pipelines that tell us whether our autonomy stack is improving or just changing.

What they're looking for

  • Currently pursuing a BS, MS, or PhD in computer science, robotics, statistics, operations research, or a related quantitative field — or bringing equivalent hands-on experience
  • Strong Python and comfort building data pipelines and analysis tooling
  • Ability to translate a fuzzy notion of "good" into a concrete, measurable definition — and to articulate why that definition is the right one
  • Solid statistical intuition: understanding variance, significance, and when a number is actually telling you something
  • Clear written and verbal communication — metric definitions have to be understood and trusted by people who didn't build them
More about this role

At Bedrock, we're moving AI out of the lab and into the real world. Our team includes veterans who helped launch Waymo, scaled Segment to a $3.2B acquisition, and grew Uber Freight to $5B in revenue. Today, we're deploying autonomous systems on heavy construction equipment across the country, improving safety on job sites and accelerating schedules on critical infrastructure projects.

We're not here debating the future of AI. We're deploying it in the real world. In just two years, we've raised $350M and achieved the first fully autonomous excavator deployments in construction.

This is where algorithms meet steel-toed boots. You'll work alongside construction veterans and world-class engineers to solve physical-world problems that simulations can't touch. If you're ready to do meaningful work on hard problems, we'd love to have you join us.

Every time Bedrock works with a new machine for tasks like grading pads, digging trenches, or loading trucks, we need a way to verify the machine is actually performing well. The Evaluation team defines what "good" looks like: building the metrics, benchmarks, and evaluation pipelines that tell us whether our autonomy stack is improving or just...

Read the full posting on Bedrock Robotics's site ↗

Engineering

Build your edge while you search

Free tools for founders and investors, plus VC Unfiltered, our take on startups, venture and the people who build them.