WitnessAI is the AI security and governance platform with network visibility, intent-based controls & runtime defense to secure every employee, model, application & agent. Backed by GV.
About the role
As a Machine Learning Engineer, you’ll design, build, and evaluate language models that power our AI security products. You’ll own the end-to-end pipeline — from dataset curation and preprocessing to experiment design, evaluation, and visualization of results. This role blends engineering and applied research, with an emphasis on producing reliable, interpretable, and safe language models. Build scalable pipelines to collect, preprocess, and manage datasets for training and evaluation of LLMs.
What they're looking for
- Experience: 2–5+ years working in machine learning or data science, ideally in a security or infrastructure-heavy environment
- Strong software engineering background (Python, testing frameworks like pytest/unittest, CI/CD tools)
- Proficiency in ML frameworks such as PyTorch
- Experience with data engineering tools (e.g., Spark, Kafka, Airflow)
- Familiarity with deploying models on cloud platforms (AWS, GCP, or Azure) and containerized environments (Docker, Kubernetes)
- Strong knowledge of ML fundamentals (supervised/unsupervised learning, deep learning, NLP)
More about this role
Job Title: Machine Learning Engineer
Type: Full-time
Team: Machine Learning
Witness AI invented intent-based AI security. While legacy tools monitor what users say to AI, we understand what they're trying to accomplish - stopping jailbreaks, data exfiltration, and shadow AI before damage occurs. We provide visibility into how employees and systems use AI - capturing prompts, responses, and agent activity - so security teams can monitor risk, investigate incidents, and enforce guardrails in real time.
As a Machine Learning Engineer, you’ll design, build, and evaluate language models that power our AI security products. You’ll own the end-to-end pipeline — from dataset curation and preprocessing to experiment design, evaluation, and visualization of results. This role blends engineering and applied research, with an emphasis on producing reliable, interpretable, and safe language models.
Build scalable pipelines to collect, preprocess, and manage datasets for training and evaluation of LLMs.
Design and run experiments to evaluate LLMs on accuracy, robustness, fairness, and safety.
Create dashboards, reports, and visualizations to communicate evaluation results, trends, and failure...
Browse similar: AI jobs · AI startup jobs · Startup jobs · San Francisco Bay Area