Startups · AI

Member of Technical Staff, Applied Research

LlamaIndex · San Francisco · Remote

← All jobs
About LlamaIndex

LlamaParse is the world. Backed by Greylock and Norwest.

About the role

We are looking for an AI Research Engineer to join our document understanding team. This role is ideal for someone who sits between applied research and strong engineering. You will work on vision-language models, document processing, data curation, synthetic data generation, benchmarking, training, fine-tuning, and post-training. The goal is simple: make our document AI systems more accurate, faster, and more cost-effective in production.

What they're looking for

  • 3–7 years of experience in machine learning engineering, applied research, or research engineering
  • Strong ML foundation, including hands-on experience benchmarking and training models
  • Strong Python skills and comfort with modern ML tooling, especially PyTorch
  • Experience with computer vision, vision-language models, NLP, document AI, OCR, extraction, or agentic AI systems
  • Ability to build experiments, evaluate results, and iterate quickly toward measurable performance improvements
  • Strong engineering judgment and ability to write clean, production-quality code
More about this role

Join us and help shape the future of AI by defining the narrative around document understanding.

We are looking for an AI Research Engineer to join our document understanding team.

This role is ideal for someone who sits between applied research and strong engineering. You will work on vision-language models, document processing, data curation, synthetic data generation, benchmarking, training, fine-tuning, and post-training. The goal is simple: make our document AI systems more accurate, faster, and more cost-effective in production.

You should be excited by frontier AI work, but equally motivated by practical product impact. This is not a pure research role where ideas stay in papers. You will be expected to prototype quickly, evaluate rigorously, and help turn promising approaches into production systems used by customers.

Develop and train vision-language models for document processing and document understanding.

Build data pipelines for data curation, synthetic data generation, labeling, and benchmark creation.

Evaluate base models and perform post-training or fine-tuning to hit specific performance targets.

Improve model accuracy, latency, and cost-effectiveness across...

Read the full posting on LlamaIndex's site ↗

Engineering

Build your edge while you search

Free tools for founders and investors, plus VC Unfiltered, our take on startups, venture and the people who build them.