Startups · AI

Sr. Software Engineer - Ingestion Core team

Databricks · San Francisco, California · On-site

← All jobs
About Databricks

Backed by Battery, Insight and Kleiner Perkins.

About the role

To enable all of this on Databricks, making data ingestion seamless is crucial. That’s the mission of the Ingestion Core Team: to make the ingestion of all data—structured and unstructured—simple, reliable, and efficient. Simplifying the complex is hard, and that’s where you come in.

What they're looking for

  • Build distributed infrastructure to ingest data from diverse sources and support streaming ingestion, incremental processing, and replication. This isn’t just about building plugin connectors
  • Reduce end-to-end latency, increase throughput, and reduce costs from the time data appears in source systems to when it is available in Delta Lake
  • Design and optimize streaming and distributed workloads for throughput, cost, latency, reliability, and scale
  • Optimize streaming workloads by exploring and applying ML techniques
  • Build monitoring and observability capabilities (customer-facing and internal) that provide visibility into ingestion workflows and the systems running them
  • Collaborate with partner teams to enable use cases like RAG and AI agents
More about this role

Deeply understanding what’s in the enterprise data has been a challenge that Databricks has been addressing by providing analytics and machine learning tools. From data warehousing with Databricks SQL to large-scale distributed processing with Spark and advanced ML tools for experimentation and model serving, we empower our customers to gain insights and drive innovation.

To enable all of this on Databricks, making data ingestion seamless is crucial. That’s the mission of the Ingestion Core Team: to make the ingestion of all data—structured and unstructured—simple, reliable, and efficient. Simplifying the complex is hard, and that’s where you come in. This role requires building distributed platform systems to incrementally ingest high-volume, petabyte-scale data from diverse sources—including cloud storage (SQS, ADLS, GCS), databases (Oracle, SQL Server, MySQL, Postgres), and file sources (Google Drive, SharePoint)—at high throughput and low cost. The data includes structured formats (JSON, Parquet, CSV) as well as unstructured data (text, images, docs, PPTs, and blobs), all of which land in Delta Lake with schema evolution and change data capture (CDC) capabilities.

Join us in...

Read the full posting on Databricks's site ↗

Engineering - Pipeline

Build your edge while you search

Free tools for founders and investors, plus VC Unfiltered, our take on startups, venture and the people who build them.