Give every agent a cloud desktop. Backed by Y Combinator.
About the role
Cua is building the infrastructure that enables general-purpose AI agents to safely and scalably use real computers and applications. We're a small team backed by Y Combinator and top-tier investors, and our open-source tools are already used by thousands of developers. As a Research Intern , you’ll help prototype, test, and benchmark multi-modal LLM-based agents - from data pipelines to orchestration systems.
What they're looking for
- Currently a PhD student in Computer Science or related field (strong Master’s considered)
- Experience in applied research with a solid publication record
- Familiarity with modern multi-modal or reasoning agents (e.g., OS-Atlas, Qwen, GUI-R1)
- Hands-on experience with PyTorch , Python , and cloud compute (AWS, GCP, etc.)
- Comfortable designing experiments, evaluating models, and working with multi-modal data
- Excited by generative AI, agent systems, and pushing the boundaries of what’s possible
More about this role
Cua is building the infrastructure that enables general-purpose AI agents to safely and scalably use real computers and applications.
We're a small team backed by Y Combinator and top-tier investors, and our open-source tools are already used by thousands of developers. As a Research Intern , you’ll help prototype, test, and benchmark multi-modal LLM-based agents - from data pipelines to orchestration systems.
You’ll collaborate with engineers and researchers to turn cutting-edge ideas into real systems and benchmarks that can be shared with the community. This is a chance to contribute to open-source research, design experiments, and explore the frontiers of agentic AI.
- Generate and curate large-scale, high-quality multi-modal data (GUIs, browsers, system UIs)
- Design and test single- and multi-agent systems for data and computer use
- Automate benchmarking of agent orchestration (with or without human-in-the-loop)
- Explore new training and inference techniques to boost reasoning and action-taking (e.g., RL-based agents)
- Develop benchmarks, tools, and datasets to evaluate agentic capabilities on Cua
- Collaborate with the founding team and contribute to research...
Browse similar: AI jobs · Startup internships · AI startup jobs · Startup jobs · Remote jobs