Startups · AI

Machine Learning Engineer

Witness AI · Bay Area · On-site

← All jobs
About Witness AI

WitnessAI is the AI security and governance platform with network visibility, intent-based controls & runtime defense to secure every employee, model, application & agent. Backed by GV.

About the role

As a Machine Learning Engineer, you’ll design, build, and evaluate language models that power our AI security products. You’ll own the end-to-end pipeline — from dataset curation and preprocessing to experiment design, evaluation, and visualization of results. This role blends engineering and applied research, with an emphasis on producing reliable, interpretable, and safe language models. Build scalable pipelines to collect, preprocess, and manage datasets for training and evaluation of LLMs.

What they're looking for

  • Experience: 2–5+ years working in machine learning or data science, ideally in a security or infrastructure-heavy environment
  • Strong software engineering background (Python, testing frameworks like pytest/unittest, CI/CD tools)
  • Proficiency in ML frameworks such as PyTorch
  • Experience with data engineering tools (e.g., Spark, Kafka, Airflow)
  • Familiarity with deploying models on cloud platforms (AWS, GCP, or Azure) and containerized environments (Docker, Kubernetes)
  • Strong knowledge of ML fundamentals (supervised/unsupervised learning, deep learning, NLP)
More about this role

Job Title: Machine Learning Engineer

Type: Full-time

Team: Machine Learning

Witness AI invented intent-based AI security. While legacy tools monitor what users say to AI, we understand what they're trying to accomplish - stopping jailbreaks, data exfiltration, and shadow AI before damage occurs. We provide visibility into how employees and systems use AI - capturing prompts, responses, and agent activity - so security teams can monitor risk, investigate incidents, and enforce guardrails in real time.

As a Machine Learning Engineer, you’ll design, build, and evaluate language models that power our AI security products. You’ll own the end-to-end pipeline — from dataset curation and preprocessing to experiment design, evaluation, and visualization of results. This role blends engineering and applied research, with an emphasis on producing reliable, interpretable, and safe language models.

Build scalable pipelines to collect, preprocess, and manage datasets for training and evaluation of LLMs.

Design and run experiments to evaluate LLMs on accuracy, robustness, fairness, and safety.

Create dashboards, reports, and visualizations to communicate evaluation results, trends, and failure...

Read the full posting on Witness AI's site ↗

Engineering

Build your edge while you search

Free tools for founders and investors, plus VC Unfiltered, our take on startups, venture and the people who build them.