Senior Site Reliability Engineer (SRE & AI Platform Operations)
ContractHybridIn English
About the job
Helloprint is seeking a Senior Site Reliability Engineer to join their hybrid team in Rotterdam. In this role, you will ensure the reliability and scalability of high-traffic distributed systems and AI platform operations. You will drive observability, automate deployments, and optimize cloud costs while leading incident response.
What you will do
- Define and enforce SLOs, SLIs, and error-budget policies
- Expand distributed observability and telemetry across microservices
- Evolve CI/CD pipelines with canary traffic shifting and automated rollbacks
- Architect and scale runtime infrastructure for AI agents and semantic pipelines
- Drive FinOps practices and optimize cloud infrastructure costs
- Lead incident response and blameless post-mortems
What you bring
- Experience operating and scaling high-traffic distributed production systems
- Strong troubleshooting skills across Linux, containers, databases, and cloud networks
- Extensive hands-on experience with Google Cloud Platform, Cloud Run, and Terraform
- Ability to monitor and optimize cloud and AI runtime costs
- Practical experience with Sentry and Google Cloud Monitoring
- Proficiency in Python or TypeScript/JavaScript and familiarity with PHP in Laravel
Nice to have
- Familiarity with operationalizing LLM integrations and vector workflows
What you get
- Opportunities for rapid growth within role and across company
- International work environment with diverse team in Rotterdam or Valencia
- Real impact in a high-growth, AI-driven company
- Ownership from day one with freedom to innovate
- Hello-Benefits including gym access, sports discount, events, and healthy meals
Summary by Personeel.com. The full job description and the application form are on Helloprint's website (helloprint.recruitee.com); Apply takes you straight there.