SURF logo

Data Architect

SURF

Groningen
Full-time
5-10 years experience
Hybrid

€4,791 - €6,845 per year

Key Skills

Data Architecture
Data Governance
Data Security
Access Control
GNU/Linux
Python
Relational and Non-Relational Databases
S3-Compatible Object Storage
Parallel File Systems
ML/AI Workflows
MLflow
Weights & Biases
Data Versioning and Lineage
Containerized Data Pipelines
GDPR and AI Act Compliance
High-Performance Computing

Job Description

Are you ready to design and build the data backbone of the Dutch AI factory, where large-scale data, AI and high-performance infrastructure come together? At SURF, you will play a key role in shaping the data architecture that enables cutting-edge research and innovation across the Netherlands. As a Data Architect, you will ensure that complex, multi-source input datasets as well as the models and artefacts are stored reliably and made accessible. Does building the data infrastructure behind Europe’s AI ambitions sounds like your kind of challenge? Apply now! Where you will work SURF is the ICT cooperative for Dutch educational and research institutions. Together with them, we work on digital services and complex innovation challenges to enhance the quality of education and research. Working at SURF means being part of a unique and open organization. You can see this in everything: the organizational structure, the composition of the project teams, the culture in our offices, and the atmosphere among colleagues. SURF offers excellent employment conditions and takes a flexible approach to work–life balance. Employees enjoy working independently, and everyone is given the space and freedom to apply and develop their talents as effectively and broadly as possible. The team you will join You will join SURF’s High Performance Machine Learning team within the Advanced Solutions for Research unit. Your colleagues work on projects such as OpenEuroLLM and GPT-NL and support researchers in making optimal use of the Dutch national supercomputer Snellius for AI workloads. The team has an open and collaborative culture, with a strong focus on sharing knowledge and supporting each other. SURF is one of the consortium partners behind the Dutch AI Factory, together with TNO, AIC4NL and Samenwerking Noord. Within this partnership, SURF is responsible for the technical infrastructure and contributes its expertise in supercomputing, providing secure access to the computing power and data needed to develop and train AI models. As part of SURF, you will therefore work directly on the development and delivery of the Dutch AI Factory. You will work closely with AI consultants, platform engineers and other experts from the consortium partners. Together, you will provide AI services to research organisations, public organisations and private-sector partners in the Netherlands and Europe. The Dutch AI Factory is also part of a wider European network of AI Factories, contributing to Europe’s ambition to strengthen its AI capabilities and technological sovereignty. Colleagues from the different consortium partners working on the AI Factory will primarily be based together in Groningen, while maintaining close connections with their respective organisations. What you will do As a Data Architect you are responsible for the design and building of the data backbone and the data portfolio of the Dutch AI factory. The data services within data portfolio ensure that the data (training data, models and artefacts) and its management can support the users’ pipelines in a secure, performant and scalable way backed by solid governance, compliance and user friendliness. Your responsibilities include: Technical development of the data portfolio on the AI factory infrastructure (data landing zones, data transfer across different storage backends, supporting pipelines for distributed data loading for multi-GPU/multi-node training, etc.) Activities to design the data governance and access control model and its technical implementation Testing and implementation of technical services to manage training data as well as model outputs (provenance, versioning, metadata, etc.) Strict security and compliance for sensitive data (e.g., via implementation of secure enclaves) Design and support building of scalable data pipelines for large-scale AI workloads Supporting integration with European data spaces and national sectoral data infrastructures Contributing to the data service catalogue and user documentation Advising project teams on their end-to-end AI workflows Your skills and experience You are an experienced Data Architect with a 5+ years of technical hands-on experience with focus on scalable, future-proof data solutions. Your leadership skills as well as extensive hands-on experience brings structure to complexity, you can manage stakeholders at various layers in organizations, and you can steer AI consultants and platform engineers to build the data backbone. You have: BSc/MSc level in computer science (including AI), data science, information science or equivalent A thorough understanding of host and file systems security, authentication and authorization principles in GNU/Linux environments Knowledge of databases (relational, non-relational), object storage (S3-compatible), shared parallel file systems (POSIX), and data catalogues Hands-on Experience with tools for full lifecyle of ML/AI workflows (e.g., MLFlow, wandb, etc.) and with data versioning and lineage tooling (e.g. DVC, LakeFS, Delta/Iceberg) Good command of Python; experience with containerised data pipelines Strong command of English; Dutch is a plus Strong hands-on experience with AI workflows and engineering tooling Experience with data governance and/or privacy regulation (e.g., GDPR, AI act) Strong advantages Experience in management of large AI datasets, with versioning, reproducibility, and automated data-quality monitoring Knowledge of NLP data pipelines (language detection, PII detection, text normalisation) Familiarity with workflow orchestration tools (e.g. Airflow) Knowledge of linked data, RDF, or SPARQL Experience with distributed data processing on compute infrastructure (HPC systems, Kubernetes) Experience with ETL/ELT pipelines, data lakes, data warehousing Knowledge of FAIR principles and experience with metadata standards What we offer A varied and challenging position for 32–40 hours per week (0.8–1.0 FTE) in a casual and collaborative organization with high standards; Extensive training opportunities and excellent benefits; A salary between 4,791 and 6,845 euros gross (based on a 40-hour workweek); 8.33 percent vacation pay and a fixed year-end bonus of 8.33 percent; 36 vacation days per year (based on a 40-hour workweek); A generous pension plan through PNO-media, where you contribute only 1/8 of the premium and SURF contributes 7/8; The option to work hybrid. You’ll receive a work-from-home allowance on your work-from-home days; A first-class NS Business Card; Work at a new location in Groningen: Niemeyer Building – Paterswoldseweg 43; A one-year contract, with the intention of converting this to permanent employment upon satisfactory performance. Applications will be accepted through October 25, 2026, via the application button. If you have any questions about the role, please feel free to contact Nicolas Renaud at [email protected]. If you have questions about the application process, please contact the recruitment team at [email protected]. Additional information A background screening may be part of the recruitment process. Please also note that job interviews will take place at the SURF office in Amsterdam. SURF takes pleasure in doing its recruitment itself; acquisition is therefore not appreciated.

Core Responsibilities

Design and build the Dutch AI Factory’s data backbone and services, including secure, scalable data storage, transfer, governance, access control, and pipelines for large-scale AI training. Manage training data and model outputs through provenance, versioning, and metadata; support compliance, integrations with data infrastructures, documentation, and advice on end-to-end AI workflows.

Requirements

Requires a bachelor’s or master’s-level background in computer science, data science, information science, or an equivalent field, plus at least five years of hands-on experience designing scalable data solutions. Candidates should bring strong knowledge of Linux security and authorization, storage and databases, AI workflows, Python, containerized pipelines, and data governance or privacy regulation; strong English is required, while Dutch is a plus.

Benefits

  • Training Opportunities
  • Vacation Pay
  • Year-End Bonus
  • 36 Vacation Days
  • Pension Plan
  • Hybrid Work
  • Work-From-Home Allowance
  • First-Class Public Transport Business Card

About SURF

Industry: IT Services and IT Consulting

Company size: 201-500 employees

SURF is a cooperative association of Dutch educational and research institutions in which the members combine their strengths. Within SURF, we work together to acquire or develop the best possible digital services, and to encourage knowledge sharing through continuous innovation. The members are the owners of SURF.

Added Today