AI Infrastructure Systems Engineer (Amsterdam & London)
HybridIn English
About the job
Join Together AI to build the backbone of AI infrastructure, designing and automating systems for large-scale GPU clusters. This hybrid role in Amsterdam focuses on maximizing performance and reliability across hardware and software.
What you will do
- Design and build fleet automation systems for GPU clusters
- Build AI Infrastructure Agents for automated deployment and remediation
- Develop Fleet Intelligence platforms for hardware health monitoring
- Maximize GPU availability, utilization, performance, and reliability
- Create automated validation systems for GPUs and networking
- Build internal platforms and developer tools for infrastructure management
What you bring
- 3+ years building distributed systems or infrastructure platforms
- Strong software engineering in Python, Go, or Rust
- Experience with platforms, automation systems, or developer infrastructure
- Experience with Linux, Kubernetes, Terraform, or Ansible
- Strong systems thinking across hardware and software
- Automation-first mindset
Nice to have
- GPU infrastructure, CUDA, NCCL, NVLink/NVSwitch
- InfiniBand or RoCE networking
- Bare-metal provisioning and lifecycle management
- Large-scale AI training or inference clusters
- Hardware health monitoring and predictive failure detection
- Distributed storage systems
Summary by Personeel.com. The full job description and the application form are on Together AI's website (job-boards.greenhouse.io); Apply takes you straight there.
More jobs at Together AI
- Research Intern, Model Shaping (Summer 2027)Amsterdam
- Senior Software Engineer Together Cloud InfrastructureAmsterdam
- Senior Software Engineer — Infra Agent SystemsAmsterdam
- Senior Program Manager, Data Center DeliveryThuiswerken
- Senior Network Engineer (Amsterdam)Amsterdam