
NexGen Cloud · Remote
PLATFORM OPERATIONS LEAD Location: UK (Remote) Department: Infrastructure Reporting to: Head of Infrastructure ABOUT NEXGEN CLOUD: NexGen Cloud is the compan...
Location: UK (Remote)
Department: Infrastructure
Reporting to: Head of Infrastructure
NexGen Cloud is the company behind Hyperstack, a full-stack AI cloud serving tens of thousands of customers from AI researchers to
enterprises running the world's most compute-intensive workloads. We deliver on-demand and private GPU infrastructure to teams who
treat performance as a requirement, not a feature.
We're a tight-knit, fast-moving team working at the cutting edge of AI cloud infrastructure. We practice what we preach, equipping
our people with AI at every level so we can solve harder problems, ship faster, and keep raising the bar for what enterprise GPU
infrastructure looks like.
This role exists to help NexGen Cloud scale the operational maturity of its cloud infrastructure as demand grows across regions,
customers and services.
You'll sit at the intersection of Infrastructure, DevOps, Engineering and Customer Experience, helping reduce operational load on
engineering teams through automation, tooling, runbooks and clear support processes. You'll play a key role in improving
reliability, observability, incident response and operational readiness across our platform.
This is a hands-on role for someone who enjoys building practical solutions, improving how teams work, and taking ownership of
operational outcomes in a fast-moving cloud environment.
Rather than a long checklist, here's what success in this role looks like:
Engineering teams
We're more interested in how you think and work than in a perfect CV. You'll likely bring a combination of the following:
Head over to our NexGen Cloud careers page to view current openings and follow us on LinkedIn and X to learn more about our
journey, newest releases and hear exciting news in the neocloud space.
Senior Infrastructure Engineer (OpenStack) Location: UK (Remote) Department: Infrastructure Reporting to: Head of Infrastructure ABOUT NEXGEN CLOUD: NexGen Cloud is the company behind Hyperstack, a full-stack AI cloud serving tens of thousands of customers from AI researchers to enterprises running the world's most compute-intensive workloads. We deliver on-demand and private GPU infrastructure to teams who treat performance as a requirement, not a feature. We're a tight-knit, fast-moving team working at the cutting edge of AI cloud infrastructure. We practice what we preach, equipping our people with AI at every level so we can solve harder problems, ship faster, and keep raising the bar for what enterprise GPU infrastructure looks like. THE ROLE: SENIOR INFRASTRUCTURE ENGINEER (OPENSTACK) This role exists because our platform is scaling quickly — and complexity comes with it. As we expand our OpenStack and Kubernetes environments globally, we need engineers who can take real ownership of how the platform is designed, operated, and improved. You'll have direct ownership over business-critical infrastructure that impacts performance, reliability, and customer experience. This is not a maintenance role. If you like solving hard problems, owning systems end-to-end, and seeing the impact of your work immediately — you'll enjoy this. WHAT YOU'LL BE DOING: Rather than a long checklist, here's what success in this role looks like: * Own the design, deployment, and operation of OpenStack and Kubernetes environments — ensuring platform performance, scalability, and resilience for GPU workloads * Build and improve infrastructure using infrastructure-as-code and GitOps practices, driving automation across provisioning, deployment, and operational workflows * Optimise GPU workload scheduling using Kubernetes and NVIDIA tooling, and implement monitoring, logging, and alerting to ensure platform stability * Lead incident response and drive continuous improvement of reliability across the platform * Maintain strong security controls across infrastructure and container layers — RBAC, network policies, and tenant isolation * Work closely with Platform, DevOps, AI, Product, and Support teams to align infrastructure capabilities with customer and platform requirements ABOUT YOU: We're more interested in how you think and work than in a perfect CV. You'll likely bring a combination of the following: ESSENTIAL * Exposure to GPU infrastructure, HPC, or large-scale compute environments * Proven experience operating Kubernetes at scale — ideally bare-metal or private cloud * Solid understanding of Linux, networking, and storage systems * Experience with infrastructure automation, CI/CD, and Git-based workflows * Strong ownership mindset — comfortable operating without heavy oversight and able to simplify and scale systems in a fast-moving environment NICE TO HAVE * Experience integrating Kubernetes with OpenStack * Ideally hands-on experience running OpenStack in production environments * Familiarity with advanced networking or cloud-native ecosystems * Contributions to open-source projects WHAT WE OFFER: * Competitive salary and annual discretionary bonus scheme * Employee wellbeing benefits * 25 days of holiday, plus public holidays * Flexible working arrangements (remote or hybrid, depending on role and location) * Real ownership and autonomy, with the trust to take initiative and experiment * The opportunity to make a visible, meaningful impact as we scale * Clear career progression and growth opportunities in a fast-growing company * A collaborative, international culture built on trust, transparency, and ownership * The chance to help shape NexGen Cloud's team, culture, and future alongside ambitious, mission-driven colleagues MORE INFORMATION Head over to our NexGen Cloud careers page to view current openings and follow us on LinkedIn and X to learn more about our journey, newest releases and hear exciting news in the neocloud space.
ELEKS is looking for an Infrastructure/GPU Cluster/Platform Operations Lead in Canada. Alberta-based candidates are strongly preferred (Calgary or Edmonton). Canada-based candidates will also be considered. ABOUT CLIENT Our customer is building a next-generation AI platform that enables organizations to securely develop, govern, and operationalize artificial intelligence while ensuring that sensitive data and organizational knowledge remain fully under their control. The platform combines advanced AI capabilities with enterprise-grade governance, security, and data sovereignty to support mission-critical decision-making. The solution serves government organizations and enterprise customers operating in highly regulated and security-sensitive environments, where reliability, accountability, and trust are essential. The platform supports intelligent decision-making across strategic planning, workforce intelligence, and organizational operations, helping customers leverage AI without compromising security, compliance, or control over their data.
WHAT WE NEED The Operations Lead is responsible for leading the day-to-day operations, service delivery, reliability, and continuous improvement of the Sakani and Ejar platforms. The role ensures platform availability, operational excellence, incident management, disaster recovery readiness, and effective collaboration between business, development, infrastructure, security, QA, and external vendors WHAT YOU'LL DO Operations Management * Lead and manage the operational support of Sakani and Ejar platforms. * Ensure platform availability, stability, performance, and reliability. * Oversee production and non-production environments. * Monitor operational KPIs, SLAs, and service health metrics. * Drive continuous service improvement initiatives. Incident & Problem Management * Lead Major Incident Management activities and coordinate resolution efforts. * Manage incident escalations across Development, Infrastructure, Database, Security, and Vendor teams. * Conduct Root Cause Analysis (RCA) and ensure implementation of corrective actions. * Track recurring issues and drive long-term solutions. Release & Change Management * Oversee production deployments, releases, and maintenance activities. * Ensure operational readiness before major releases. * Review implementation and rollback plans. * Ensure compliance with change management processes and governance standards. Platform & DevOps Operations * Collaborate with DevOps teams to enhance automation, reliability, and operational efficiency. * Ensure monitoring, logging, alerting, and observability capabilities are maintained. * Support Kubernetes, cloud infrastructure, CI/CD pipelines, and platform operations. * Drive operational excellence through automation and process optimization. Disaster Recovery & Business Continuity * Lead Disaster Recovery (DR) planning, testing, and execution activities. * Ensure Recovery Time Objective (RTO) and Recovery Point Objective (RPO) targets are achieved. * Maintain operational runbooks and recovery procedures. * Coordinate periodic DR drills and readiness assessments. Stakeholder Management * Act as the primary operational contact for business and technical stakeholders. * Coordinate with internal teams, external vendors, and service providers. * Prepare and present operational reports, service reviews, and executive updates. * Facilitate operational governance and service review meetings. Vendor & Service Management * Manage vendor performance against agreed Service Level Agreements (SLAs). * Coordinate operational activities with third-party providers. * Handle escalations and ensure timely issue resolution. Performance & Capacity Management * Monitor platform performance, utilization, and capacity trends. * Identify and address performance bottlenecks. * Plan future capacity requirements and scalability improvements. Security, Risk & Compliance * Ensure compliance with organizational policies, security standards, and regulatory requirements. * Support security audits, risk assessments, and compliance initiatives. * Track and mitigate operational risks. Key Performance Indicators (KPIs) * Platform Availability ≥ 99.9%. * SLA Compliance ≥ 95%. * Change Success Rate ≥ 98%. * Reduction in Mean Time to Recovery (MTTR). * Successful Disaster Recovery testing and execution. * Operational Risk Reduction. * Stakeholder and Customer Satisfaction. Reporting To IT Operations Manager / Head of Technology Operations. Scope Responsible for operational governance, service reliability, incident management, platform operations, disaster recovery readiness, and service delivery across. WHAT YOU HAVE Qualifications * Bachelor’s degree in Computer Science, Information Technology, Engineering, or a related field. * Minimum 8 years of experience in IT Operations, Service Delivery, Infrastructure, or DevOps. * Minimum 3 years of experience in a leadership or management role. * Proven experience managing large-scale enterprise platforms and critical business services. Experience working with cross-functional teams and external vendors.