# Senior ML Research Engineer, Virtual Cell at SandboxAQ

- Company: SandboxAQ
- What the company does: SandboxAQ leverages the compound effects of AI and advanced computing to address some of the biggest challenges impacting society. SandboxAQ technologies include AI simulation, cryptography management for cybersecurity, and AI sensing for global organizations.
- Company website: https://www.sandboxaq.com
- Type: Startups (AI role)
- Level: Senior
- Location: United States
- Work setup: Remote
- Pay: $134K to $252K base salary per year (USD)
- Posted: 2026-07-16
- Apply by: 2026-10-08
- Apply: https://jobs.ashbyhq.com/sandboxaq/aeb6fcd4-d665-4322-8486-916567f29fd2
- Page: https://www.1752.vc/careers/jobs/sandboxaq-senior-ml-research-engineer-virtual-cell/

## About the role

The AI Sim R&D team builds leading-edge ML and physics-based models ("LQMs") to advance drug discovery. Within this team, AQCell is our virtual cell platform: it takes a cell representation (e.g. basal gene expression) and a perturbation descriptor (e.g.

## What they're looking for

- Academic Foundation: Bachelor's degree in a scientific or quantitative field (Computer Science, Physics, Mathematics, Biology, Chemistry, or related), an advanced degree (MS or PhD) is preferred
- Large-Scale Data Management: Required experience managing large-scale datasets and managing the training of models over them, including data ingestion, cleaning, versioning, and pipeline maintenance at scale
- Software Engineering: Strong Python programming skills and experience with modern ML frameworks (e.g. PyTorch, JAX) and experiment tracking/data versioning tools (e.g. Weights & Biases)
- Scientific Rigor: Ability to design sound evaluation methodology (e.g. train/test splitting strategies, held-out generalization tests) and to critically interpret model performance against meaningful baselines
- Bioinformatics & Computational Biology: Experience with bioinformatics and computational biology data analysis tools, particularly transcriptomics harmonization and normalization tools (e.g. batch correction, pseudobulking, gene ID mapping/standardization)
- Domain Datasets: Familiarity with public perturbation or drug-sensitivity datasets such as LINCS L1000, GDSC, DepMap, or single-cell perturbation atlases

Tags: AI Simulation
