Mirrai Careers
Resume BuilderCareer Test
InsightsPricing
Get Started Free
Jobs/Infrastructure/GPU Cluster/Platform Operations Lead

Infrastructure/GPU Cluster/Platform Operations Lead

eleks

Remote (Canada) Posted 3w ago
Apply on company site
ELEKS is looking for an Infrastructure/GPU Cluster/Platform Operations Lead in Canada. Alberta-based candidates are strongly preferred (Calgary or Edmonton). Canada-based candidates will also be considered.   ABOUT CLIENT Our customer is building a next-generation AI platform that enables organizations to securely develop, govern, and operationalize artificial intelligence while ensuring that sensitive data and organizational knowledge remain fully under their control. The platform combines advanced AI capabilities with enterprise-grade governance, security, and data sovereignty to support mission-critical decision-making. The solution serves government organizations and enterprise customers operating in highly regulated and security-sensitive environments, where reliability, accountability, and trust are essential. The platform supports intelligent decision-making across strategic planning, workforce intelligence, and organizational operations, helping customers leverage AI without compromising security, compliance, or control over their data. REQUIREMENTS 8+ years of Infrastructure Engineering or Platform Operations experience Experience managing GPU clusters for AI workloads Strong Kubernetes administration skills Experience with NVIDIA GPU technologies and CUDA ecosystem Experience with cloud infrastructure (Azure, AWS or GCP) Knowledge of storage, networking, and high-performance computing environments Experience implementing Infrastructure as Code (Terraform or similar) Strong operational leadership skills Experience supporting AI platform infrastructure Upper-Intermediate or higher level of English RESPONSIBILITIES Lead GPU infrastructure design and operations Manage Kubernetes-based AI platform environments Optimize infrastructure for AI training and inference workloads Define operational standards and reliability practices Collaborate with AI engineering teams Implement monitoring, security, and disaster recovery strategies Lead infrastructure capacity planning Support technical roadmap and infrastructure evolution

See how well you match this job

Upload your resume and we’ll score your fit for this role and 6 similar roles — then tailor your CV to it with AI. Free, no credit card.

Check your match

Similar jobs

  • MLOps/LLMOps Architect

    eleks

    Remote (Canada)
  • AI/Data Platform Architect (Digital Twin/Knowledge Graph)

    eleks

    Remote (Canada)
  • Senior Infrastructure/ DevOps Engineer, Fintech

    lazer

    Remote
  • Software Engineer, GPU Infrastructure (HPC)

    Cohere

    Remote
  • Staff Cloud Operations Engineer (Toronto, Canada)

    extremenetworks

    Toronto, Canada
  • Staff AI Platform & Agent Runtime Engineer

    eqbank

    Toronto
Apply on company site

Want more roles like this? Browse fresh jobs or tailor your resume with AI.

Mirrai Careers

AI-powered career platform: build resumes, match jobs, and plan your career.

Product

  • All Tools
  • Resume Builder
  • Career Test
  • Pricing

Legal

  • Privacy Policy
  • Terms of Service
  • Fair Use Policy

Company

MIRRAI CHAT LTD (Company No. 16403306)

71-75 Shelton Street, Covent Garden

London, WC2H 9JQ, UNITED KINGDOM

[email protected]

© 2026 Mirrai Careers. All rights reserved.