Data Engineer
Key Skills
Job Description
We’re a fast-growing AI scale-up working with analysing social data to predict human behaviour at scale. We’re building agentic AI systems that need to reason over the real world and the real world produces terrible data. Billions of records. Unstructured text. Conflicting sources. Missing information. Constantly evolving datasets. And in our world, an AI agent confidently reasoning over the wrong representation of that data can be worse than having no answer at all. Following our Series A, we’re looking for a Data Engineer to join our team in the Netherlands. This isn't traditional data engineering. You willl own the path from raw open-source intelligence data to the representations our AI agents and analysts actually reason over. That means: Ingesting and transforming data at billion-row scale Building agent-ready graphical, temporal and provenance-aware data structures Designing pipelines for large-scale language model inference Making messy, incomplete and adversarial data usable Ensuring every derived fact can be traced back to its source Working across Python, PySpark, Iceberg, ClickHouse and OpenSearch You’ll work directly with AI engineers, behavioural scientists and end users. Sometimes you’ll build a POC quickly to solve an immediate operational problem. Other times you’ll turn what we learn into infrastructure that becomes part of our core platform. We’re particularly interested in engineers who have worked with large-scale, text-heavy and imperfect datasets and know what happens when data systems start to genuinely hurt at scale. Experience with OSINT, graph databases, temporal data, sovereign infrastructure or secured/ regulated environments is a plus, but not essential. What matters is engineering depth, comfort with ambiguity and an AI-native way of working. If building the data substrate that autonomous AI systems ultimately trust sounds like your kind of problem, please apply here.
Core Responsibilities
Ingest and transform large-scale, imperfect open-source intelligence data, building graph-based, temporal, and provenance-aware representations that AI agents and analysts can use. Design pipelines for large language model inference, ensure derived facts can be traced to their sources, and develop both rapid proofs of concept and durable platform infrastructure.
Requirements
The role calls for strong engineering depth, comfort with ambiguity, and an AI-native approach, particularly with large-scale, text-heavy, imperfect datasets. Experience with OSINT, graph databases, temporal data, sovereign infrastructure, or secured and regulated environments is a plus but not essential.
About Stealth Startup
Industry: Technology, Information and Internet
Company size: 11-50 employees
A network for entrepreneurs building in stealth. Submit your information here so investors can find you: harmonic.ai/get-discovered