Personeel.com

AI Infrastructure Systems Engineer (Amsterdam & London)

Amsterdam30+ days ago

HybridIn English
Apply

About the job

Join Together AI to build the backbone of AI infrastructure, designing and automating systems for large-scale GPU clusters. This hybrid role in Amsterdam focuses on maximizing performance and reliability across hardware and software.

What you will do

  • Design and build fleet automation systems for GPU clusters
  • Build AI Infrastructure Agents for automated deployment and remediation
  • Develop Fleet Intelligence platforms for hardware health monitoring
  • Maximize GPU availability, utilization, performance, and reliability
  • Create automated validation systems for GPUs and networking
  • Build internal platforms and developer tools for infrastructure management

What you bring

  • 3+ years building distributed systems or infrastructure platforms
  • Strong software engineering in Python, Go, or Rust
  • Experience with platforms, automation systems, or developer infrastructure
  • Experience with Linux, Kubernetes, Terraform, or Ansible
  • Strong systems thinking across hardware and software
  • Automation-first mindset

Nice to have

  • GPU infrastructure, CUDA, NCCL, NVLink/NVSwitch
  • InfiniBand or RoCE networking
  • Bare-metal provisioning and lifecycle management
  • Large-scale AI training or inference clusters
  • Hardware health monitoring and predictive failure detection
  • Distributed storage systems

Summary by Personeel.com. The full job description and the application form are on Together AI's website (job-boards.greenhouse.io); Apply takes you straight there.

More jobs at Together AI

All 10 jobs at Together AI
Apply