Nebul logo

Tech Lead – AI Backend

Nebul

Leiden
Full-time
10+ years experience
On-site

Key Skills

Go
Backend Engineering
Technical Leadership
Microservices
Distributed Systems
API Design
Kubernetes
Cloud-Native Environments
LLM Inference
Model Serving
GPU Orchestration
Asynchronous Processing
Fault Tolerance
System Architecture
Observability
Production Reliability

Job Description

About Nebul At Nebul, we’re building Europe’s sovereign AI cloud — secure, high-performance and purpose-built as a European alternative to the American hyperscalers. Our NeoCloud platform runs on Nebul-owned infrastructure and open-source technologies, giving us control from the hardware layer all the way through to the AI and cloud services our customers use. Within our AI engineering teams, we’re building the backend and platform capabilities needed to run modern AI workloads at scale. That means everything from LLM inference and model serving to Kubernetes-based services, GPU orchestration and distributed AI applications. We’re looking for a Tech Lead – AI Backend, Kubernetes & GPU Platform who combines strong backend engineering experience with hands-on technical leadership. What You’ll Be Doing This is a deeply technical role. You’ll help shape the architecture behind our AI platform, while staying close to the software being built. Your focus will be on designing scalable backend systems, APIs and orchestration services that allow AI workloads to run reliably across Kubernetes and GPU infrastructure. You’ll work on systems supporting LLM inference, model serving, AI applications, multi-tenant environments and distributed GPU workloads. You’ll lead architectural decisions around microservices, APIs, asynchronous workflows, state management and distributed systems, primarily using Go. You’ll work closely with AI, platform, infrastructure and networking engineers to make sure the software and infrastructure layers work together as one system. This is not an ML research role and you do not need to be a machine-learning scientist. You do, however, need to understand how modern LLMs and AI applications behave in production and what that means for scalability, availability, networking and GPU utilisation. Alongside building software yourself, you’ll guide engineers, lead design discussions, identify bottlenecks before they become scaling problems and help set the engineering standards for reliability, observability, testing and maintainability. What You Bring You have extensive backend engineering experience and strong production experience with Go / Golang — this is a must-have. You’ve worked as a Senior, Staff, Principal, Lead Engineer or Tech Lead and have designed large-scale microservices and distributed systems in production. You understand APIs, service-to-service communication, asynchronous processing, databases, queues, caching, retries, idempotency and fault tolerance. You’re comfortable working in Kubernetes and cloud-native environments and understand Kubernetes beyond simply deploying applications. Most importantly, you can take ownership of complex technical problems, create clear architectural direction and help other engineers make better technical decisions. Experience with Python, LLM inference, model serving, AI agents, RAG, GPU workloads, Kubernetes controllers, NVIDIA infrastructure, vLLM, Triton, KServe, Ray, Kafka, NATS, Redis or OpenTelemetry would be a strong advantage. This Role Is Not This is not a traditional Engineering Manager, Scrum, DevOps or ML research role. You won’t spend your days managing delivery plans, training foundation models, manually provisioning infrastructure or writing endless Terraform and YAML. Your primary responsibility is technical leadership across backend software and AI platform engineering. Eligibility You must already live and work in the Netherlands and be able to travel to our Leiden office. We offer visa sponsorship, but only for candidates who are already based in the Netherlands. English fluency is required. Dutch is not. Build the Backend Behind Europe’s AI Infrastructure Running AI in production is about much more than the model. It requires scalable backend architecture, intelligent orchestration, reliable distributed systems and infrastructure capable of handling demanding GPU workloads. If that is the kind of engineering challenge you enjoy, we’d like to hear from you. Apply through Frank Poll and help us build the technology powering Europe’s sovereign AI future.

Core Responsibilities

Design and build scalable backend systems, APIs, and orchestration services for AI workloads running across Kubernetes and GPU infrastructure. Lead architectural decisions and technical standards while guiding engineers and collaborating with AI, platform, infrastructure, and networking teams.

Requirements

Extensive production backend engineering experience with Go is required, along with experience designing large-scale microservices and distributed systems. Candidates should understand Kubernetes and cloud-native environments, production AI workloads, and core distributed-system concepts, and be able to provide technical direction and mentor engineers.

About Nebul

Industry: Technology, Information and Internet

Company size: 51-200 employees

We eliminate compliance risks and return full ownership of your data to you, offering a 100% European alternative to global hyperscalers. ✔︎ The Foundation: European NeoCloud A secure, full-stack cloud built for compliance (GDPR/AI-Act), provisioning CPU, DPU, GPU and APU's. It features 1-click deployment, managed Kubernetes, managed databases, and sovereign NVMe storage to make complex infrastructure simple. ✔︎ The Brain: Private AI Factory A controlled environment to transform data into intelligence. Access a curated pack of open-source models (Llama, Mistral) and feed autonomous agents or secure workflows via AI Studio APIs, all governed by a private AI Observer. Expanding your IP with tailored AI, instead of open-sourcing your IP with public and generic AI. All on your terms. ✔︎ The Engine: GPU SuperClusters Elite-tier power for massive scale. Leveraging NVIDIA Blackwell and Hopper architectures, it scales from a single GPU to 1,000+ nodes, fueled by memory-based storage to eliminate data bottlenecks. Nebul leverages the latest NVIDIA Supercomputers to power your AI applications from efficient green energy data centers strategically located across the European Continent. In June 2024, Nebul announced €20M Euro funding round to facilitate expansion of EU Sovereign AI Cloud & Data Centers, serving European Native AI infrastructure projects and related engineering support. This funding allows Nebul to serve its expanding customer base and AI infrastructure demand in Europe. For project inquiries: Contact us here: https://nebul.com/contact/#get-in-touch Or email us: [email protected]

Added 2 days ago