Key Skills
Job Description
Data Engineer Amsterdam | In office | Full-time | 1-3 years experience About Moonlit Moonlit is the legal publisher of today. From our headquarters in Amsterdam we serve tens of thousands of legal professionals globally, either through our platform, our API or our MCP server. Behind it all sits Europe's largest legal database, enriched and structured to power semantic search, AI assistants and the legal AI products built on top of us. Our mission Make legal information more efficient, transparent and accessible. For everyone. The role As a Data Engineer at Moonlit you'll build the data layer that powers a platform used daily by lawyers, judges and policymakers. Our users work at the highest levels and demand quality. You'll design and implement the pipelines and infrastructure that bring millions of legal documents into our ecosystem, keep them accurate and up to date and make them fast to search. We build enterprise and government grade software. What you will do Build and optimize ETL pipelines using Python, PySpark and Databricks Add new legal data sources: scraping, parsing and normalizing documents in formats like HTML, XML and PDF Enrich and structure existing datasets (metadata, references between documents, versioning of legislation) Maintain and optimize our Elasticsearch indexes for fast and reliable search Monitor data quality, freshness and coverage across all sources Improve performance, reliability and cost of existing pipelines Manage data infrastructure on Azure Collaborate with platform developers, AI engineers, legal experts and business stakeholders Data ownership Own end-to-end data workflows: from source to searchable, enriched document Build for production: tested, monitored and scalable Be the go-to person for how our data is structured and where it comes from Our tech stack Data processing: Python, PySpark, Databricks Search & storage: Elasticsearch, Turbopuffer, Azure SQL Cloud: Azure, AWS, GCP Who are you? 1-3 years of experience in data engineering or a related field Strong Python skills with experience in PySpark and large-scale data processing Solid SQL and a good understanding of data modeling Hands-on experience with at least one major cloud platform (Azure, AWS or GCP) Experience building ETL pipelines and data transformation workflows Comfortable with messy, unstructured data Experience with Databricks and/or Elasticsearch is a plus A pragmatic mindset: you choose robust solutions over clever ones Why Moonlit? At Moonlit your work has real-world impact. The data systems you build support tens of thousands of legal professionals globally and have the potential to help millions gain better access to justice. We're a quality-first, engineering-driven team that chooses scalable, modern and elegant technologies. You'll have room to take initiative, shape how we handle data and grow personally and professionally in a fast-paced scale-up. You'll work with talented developers and legal experts in an open culture. What we offer Competitive salary with equity package A very nice office with a big garden An international, mission-driven work environment
Core Responsibilities
Build and maintain scalable ETL pipelines and data infrastructure to ingest, normalize, enrich, and update legal documents from multiple sources. Optimize search indexes and production workflows while monitoring data quality, freshness, performance, reliability, and cost.
Requirements
Requires 1–3 years of data engineering or related experience, strong Python skills, experience with PySpark and large-scale data processing, and solid SQL and data modeling knowledge. Candidates should have built ETL workflows, used at least one major cloud platform, and be comfortable working with messy, unstructured data; Databricks or Elasticsearch experience is a plus.
Benefits
- Equity Package
- Office With a Garden
- International, Mission-Driven Work Environment
About Moonlit
Industry: Software Development
Company size: 11-50 employees
Moonlit continuously ingests hundreds of official sources across jurisdictions and applies millions of enrichments to transform raw publications into structured, connected legal intelligence. Get direct access through our platform, the MCP or the API-connection.