
Manager, Cloud Infrastructure & DevOps
Key Skills
Job Description
Position Summary
Thales delivers cloud-based authentication and identity services to high-profile customers who depend on uncompromising security and availability. We are expanding our leadership team to scale and continuously improve IAM services for North American customers.
You will lead a team responsible for cloud infrastructure, automation, and service reliability. You will advance DevOps, SRE, and AIOps practices to deliver secure, scalable services across AWS/GCP, partnering closely with Engineering, Security, Service Desk, and customer-facing teams.
Key Responsibilities & Leadership Impact
Leadership & Team Management
Lead, coach, and grow a team of Cloud/DevOps engineers; drive hiring, onboarding, performance management, and career development.
Set and track OKRs, manage priorities, and align execution with IAM Services objectives.
Build a culture of ownership and operational excellence, partnering with global engineering leaders on standards and the long-term cloud roadmap.
Cloud Infrastructure Ownership
Own end-to-end operations of the IAM cloud platform across AWS and GCP, ensuring resilience, scalability, performance, and security.
Drive modernization with cloud-native architectures, container orchestration, platform engineering, and zero-trust principles.
Lead roadmap planning, capacity management, and cost/efficiency optimization to meet growth and SLA commitments.
DevOps, Automation & Platform Engineering
Lead the design and governance of CI/CD, Infrastructure as Code (IaC), and automation frameworks that enable safe, repeatable change.
Own implementation and lifecycle of tooling such as Terraform, Helm, Ansible, and GitLab CI/CD.
Enforce secure engineering standards and reusable patterns through reviews, automation, and continuous improvement.
Establish best practices for observability, incident response, and SRE (SLIs/SLOs, error budgets, post-incident reviews).
AI, AIOps & Intelligent Operations
Lead adoption of AIOps to improve MTTR and reliability through anomaly detection, noise reduction, predictive capacity planning, and automated remediation.
Partner with architects to support AI workloads and enable LLM-assisted operations (runbooks, incident summaries, knowledge search) with secure data access and scalable infrastructure.
Champion responsible AI with clear governance, auditability, and secure-by-design practices aligned to Thales and customer requirements.
Operational Excellence
Ensure 24/7 availability of critical IAM services through strong L2/L3 support, operational readiness, and measurable outcomes.
Own on-call and escalation frameworks; continuously improve reliability, change management, and release safety.
Lead incident management end-to-end: coordination, stakeholder communication, root cause analysis, and systemic corrective actions.
Cross-Functional & Customer Engagement
Collaborate with R&D, Security, Customer Success, Sales Engineering, and Service Desk to deliver reliable outcomes for customers and complex deployments.
Join customer engagements as needed (architecture reviews, escalations) to represent Operations leadership and reinforce customer trust.
Represent Cloud Operations in strategic planning, audits, risk management, and executive communications.
Minimum Qualifications
People leadership experience managing and developing technical teams (coaching, feedback, delivery accountability).
5+ years deploying and supporting large-scale cloud applications in production.
5+ years with infrastructure automation and DevOps tooling including Terraform, Helm/Ansible, GitLab pipelines, and Python.
Deep expertise with Kubernetes (e.g., K8s, OpenShift) in production.
Strong understanding of change management, SLA/uptime commitments, operational playbooks, and escalation workflows.
Experience operating production systems at scale in AWS and/or GCP.
Strong knowledge of distributed systems (REST APIs, microservices) and data stores (relational/NoSQL).
Strong Linux/Windows fundamentals and solid understanding of cloud security best practices.
Excellent communication and ability to lead in a geographically distributed environment.
Preferred Qualifications
Bachelor’s degree in Computer Science, Engineering, or related field.
Experience with AIOps, observability, and SRE and/or platform engineering practices.
Interested?
Apply now. We look forward to receiving your application.
#LI-EC1
Core Responsibilities
You will lead a team of engineers to manage cloud infrastructure, automation, and service reliability across AWS and GCP platforms. The role involves driving DevOps and AIOps practices while ensuring 24/7 availability and security for critical IAM services.
Requirements
The position requires at least 5 years of experience in large-scale cloud application support and infrastructure automation. Candidates must possess deep expertise in Kubernetes, DevOps tooling, and proven experience in people leadership.
About Thales
Industry: Defense and Space Manufacturing
Company size: 10,001+ employees
Thales (Euronext Paris: HO) is a global leader in advanced technologies for the Defence, Aerospace, and Cyber & Digital sectors. Its portfolio of innovative products and services addresses several major challenges: sovereignty, security, sustainability and inclusion. The Group invests more than €4 billion per year in Research & Development in key areas, particularly for critical environments, such as Artificial Intelligence, cybersecurity, quantum and cloud technologies. Thales has more than 83,000 employees in 68 countries. In 2024, the Group generated sales of €20.6 billion.