Stealth Startup logo

Data Engineer

Stealth Startup

Amsterdam
Full-time
5-10 years experience
On-site

Key Skills

Data Engineering
Python
PySpark
Apache Iceberg
ClickHouse
OpenSearch
Data Ingestion
Data Transformation
Large-Scale Data Processing
Graph Data Structures
Temporal Data
Data Provenance
Large Language Model Inference
Unstructured Data
Pipeline Design
AI Systems

Job Description

We’re a fast-growing AI scale-up working with analysing social data to predict human behaviour at scale. We’re building agentic AI systems that need to reason over the real world and the real world produces terrible data. Billions of records. Unstructured text. Conflicting sources. Missing information. Constantly evolving datasets. And in our world, an AI agent confidently reasoning over the wrong representation of that data can be worse than having no answer at all. Following our Series A, we’re looking for a Data Engineer to join our team in the Netherlands. This isn't traditional data engineering. You willl own the path from raw open-source intelligence data to the representations our AI agents and analysts actually reason over. That means: Ingesting and transforming data at billion-row scale Building agent-ready graphical, temporal and provenance-aware data structures Designing pipelines for large-scale language model inference Making messy, incomplete and adversarial data usable Ensuring every derived fact can be traced back to its source Working across Python, PySpark, Iceberg, ClickHouse and OpenSearch You’ll work directly with AI engineers, behavioural scientists and end users. Sometimes you’ll build a POC quickly to solve an immediate operational problem. Other times you’ll turn what we learn into infrastructure that becomes part of our core platform. We’re particularly interested in engineers who have worked with large-scale, text-heavy and imperfect datasets and know what happens when data systems start to genuinely hurt at scale. Experience with OSINT, graph databases, temporal data, sovereign infrastructure or secured/ regulated environments is a plus, but not essential. What matters is engineering depth, comfort with ambiguity and an AI-native way of working. If building the data substrate that autonomous AI systems ultimately trust sounds like your kind of problem, please apply here.

Core Responsibilities

Ingest and transform large-scale, imperfect open-source intelligence data, building graph-based, temporal, and provenance-aware representations that AI agents and analysts can use. Design pipelines for large language model inference, ensure derived facts can be traced to their sources, and develop both rapid proofs of concept and durable platform infrastructure.

Requirements

The role calls for strong engineering depth, comfort with ambiguity, and an AI-native approach, particularly with large-scale, text-heavy, imperfect datasets. Experience with OSINT, graph databases, temporal data, sovereign infrastructure, or secured and regulated environments is a plus but not essential.

About Stealth Startup

Industry: Technology, Information and Internet

Company size: 11-50 employees

A network for entrepreneurs building in stealth. Submit your information here so investors can find you: harmonic.ai/get-discovered

Added 2 days ago