Empower your IT staff & propel your digital transformation with ScienceLogic's AI-driven observability and IT infrastructure monitoring platform. Backed by NEA.
About the role
Design and own evaluation harnesses for LLM and agentic outputs — golden sets, regression suites, and rubric-based scoring. Build and calibrate LLM-as-judge pipelines; validate judges against human labels and control for their bias and variance.
What they're looking for
- Bachelor's or Master's in Data Science, Computer Science, Statistics, Mathematics, or a related field or equivalent experience
- US Citizen or Green Card Holder
- 3+ years in data science, ML, or applied quantitative analysis
- Strong applied statistics, with the judgment to design sound experiments and significance tests on noisy, non-deterministic outputs (not just clean A/B conversion)
- Experience building, deploying, and monitoring predictive or time-series models in production: forecasting, anomaly detection, or trend analysis, including recalibration as data shifts
- Demonstrated work evaluating, analyzing, or improving LLM or NLP systems: eval design, quality measurement, retrieval evaluation, or agent analysis
More about this role
ScienceLogic is redefining IT operations for the modern enterprise. Our AIOps platform empowers organizations to achieve Autonomic IT — where systems are self-healing, self-optimizing, and seamlessly aligned with business outcomes. We help enterprises and service providers gain unified visibility across hybrid and multi-cloud environments, automate workflows, and unlock performance at scale.
We’re accelerating digital transformation through the power of automation, AI, and analytics — giving IT and business leaders the tools to deliver superior customer experiences, drive efficiency, and innovate with confidence.
We're looking for a strong Data Scientist to join our growing Data Science team. We run a suite of small, locally-hosted language models in production — not a single frontier API. That deliberate architecture defines this role: each model is more constrained than a giant hosted one, so product quality comes from how well we evaluate, route, prompt, ground, and orchestrate the models we have. Your job is to get the best possible outcomes out of that suite.
This is not classical predictive modeling. The object of measurement is the LLM system itself — its answers,...
Browse similar: Startup jobs · Remote jobs