
ELEKS · Remote (Canada)
ELEKS is looking for an Infrastructure/GPU Cluster/Platform Operations Lead in Canada. Alberta-based candidates are strongly preferred (Calgary or Edmonton). Ca...
ELEKS is looking for an Infrastructure/GPU Cluster/Platform Operations Lead in Canada.
Alberta-based candidates are strongly preferred (Calgary or Edmonton). Canada-based candidates will also be considered.
Our customer is building a next-generation AI platform that enables organizations to securely develop, govern, and operationalize artificial intelligence while ensuring that sensitive data and organizational knowledge remain fully under their control. The platform combines advanced AI capabilities with enterprise-grade governance, security, and data sovereignty to support mission-critical decision-making.
The solution serves government organizations and enterprise customers operating in highly regulated and security-sensitive environments, where reliability, accountability, and trust are essential. The platform supports intelligent decision-making across strategic planning, workforce intelligence, and organizational operations, helping customers leverage AI without compromising security, compliance, or control over their data.
🎤 WHY VOIZE? BECAUSE WE’RE MORE THAN JUST A JOB! At voize, we believe the greatest gift to frontline workers is time - time to care, connect, and be present. Today, that time is lost to busywork and complex systems that pull them away from what matters most: people. Our vision is to change that by building AI companions that seamlessly take over digital workflows. We don't replace humans with technology - we amplify their impact. Our mission is backed with a $50M Series A funding led by Balderton Capital, with support from HV Capital, Y Combinator and other leading VCs. Today, 2,000+ facilities trust voize, and over 200,000 users rely on our AI companion to ease their daily workload. As a dynamic team, we combine first-in-class technology with meaningful social impact. And now, we’re looking for you to join us on this mission! 💡 YOUR MISSION: DELIGHT CUSTOMERS, OPTIMIZE PROCESSES! As DevOps Engineer, you will build and own voize's platform and infrastructure — from cloud and edge systems to ML data pipelines; ensuring our systems are secure, reliable, and compliant. You will enable product and ML teams to ship state-of-the-art healthcare AI with confidence to hospitals, care facilities, and mobile devices across Europe. 🚀 YOUR DAILY BUSINESS - NO TWO DAYS ARE ALIKE * Own and operate our Kubernetes clusters (AWS EKS production, bare-metal K3s for ML training, on-premises appliances running K3s) using GitOps (FluxCD) and Infrastructure as Code (CloudFormation, Ansible) * Manage the lifecycle of on-premises gateway appliances deployed at customer sites — VM image builds, TLS certificate automation, staged rollouts via GitOps, remote monitoring and troubleshooting * Build and maintain monitoring, alerting, and observability infrastructure (Prometheus, Grafana, Loki, Tempo, OpenTelemetry) across cloud and edge environments * Drive compliance automation — security hardening, access controls, audit logging, encrypted secrets management (SOPS/KMS), and evidence collection for C5, HIPAA, and HDS certifications * Support and scale ML training and data processing infrastructure — GPU cluster management, training job orchestration, and data pipeline reliability, working closely with the ML team to power state-of-the-art speech and language models * Incident response — detection, triage, resolution, and post-incident reviews for infrastructure issues. 🤝 YOUR SKILLSET - WHAT YOU BRING TO THE TABLE * Strong experience operating Kubernetes in production (cluster operations, troubleshooting, networking, storage) * Experience with GitOps workflows * Experience working with monitoring and observability stacks * Security-minded — you have strong understanding of encryption, access controls, secrets management * Solid Linux systems administration skills, and experience with CI/CD pipeline design and maintenance. * You are comfortable with taking shifts in on-call rotations. We typically share a time zone with our customers (currently CET, with potential growth to EST). 🌱 GROWING TOGETHER – WHAT YOU CAN EXPECT AT VOIZE * We are a fast-growing startup, that means you will tackle challenges, grow quickly, and make a real impact, giving frontline workers more time for people * We foster an open, collaborative culture with regular team events, whether you work from our Berlin office or remotely across Germany * Become a co-creator of our success with stock options * Generous perks: 30 vacation days plus your birthday off, Germany Transport Ticket, Urban Sports Club, regular company off-sites, and access to learning platforms such as Blinkist and Audible, plus free language courses * You decide when you work best, that means flexible working hours and a good hybrid set-up. ✨ READY TO TALK? APPLY NOW! 🚀 We look forward to your application and can’t wait to meet you – no matter who you are or what background you have!
Company Introduction We exist to wow our customers. We know we’re doing the right thing when we hear our customers say, “How did we ever live without Coupang?” Born out of an obsession to make shopping, eating, and living easier than ever, we’re collectively disrupting the multi-billion-dollar e-commerce industry from the ground up. We are one of the fastest-growing e-commerce companies that established an unparalleled reputation for being a dominant and reliable force in South Korean commerce. We are proud to have the best of both worlds — a startup culture with the resources of a large global public company. This fuels us to continue our growth and launch new services at the speed we have been since our inception. We are all entrepreneurs surrounded by opportunities to drive new initiatives and innovations. At our core, we are bold and ambitious people that like to get our hands dirty and make a hands-on impact. At Coupang, you will see yourself, your colleagues, your team, and the company grow every day. Our mission to build the future of commerce is real. We push the boundaries of what’s possible to solve problems and break traditional tradeoffs. Join Coupang now to create an epic experience in this always-on, high-tech, and hyper-connected world. About Us Coupang is at the forefront of the AI and high-performance computing (HPC) revolution. We are building a next-generation cloud platform designed to provide developers, researchers, and enterprises with seamless, scalable, and powerful access to accelerated computing. As the demand for AI, machine learning, and data-intensive workloads skyrockets, we are looking for a visionary product leader to define the future of our core compute offerings. Job Overview CIC is looking for a Group Product Manager to own the foundational compute platform that powers enterprise AI workloads. This role spans fleet management, capacity management, bare metal, virtualized compute, node lifecycle, placement, reservations, and infrastructure-level customer experience. The Product Manager will be responsible for turning physical GPU and CPU capacity into reliable, customer-ready, observable, and billable compute products. This includes defining how capacity is reserved, provisioned, validated, monitored, maintained, packaged, and exposed to enterprise customers. This role sits at the compute foundation layer. It enables higher-level orchestration and workload services such as Kubernetes, Slurm, Ray, jobs, notebooks, and inference endpoints. The candidate should understand how those orchestration and workload systems depend on foundational compute infrastructure. Key Responsibilities * Define the product strategy and roadmap for CIC's compute platform across fleet, capacity, bare metal, and virtualized compute, working backwards from customer AI workload requirements. * Own the product lifecycle from infrastructure capacity to customer-ready compute, including reservation, provisioning, lifecycle actions, observability, billing integration, maintenance, and deprecation, with clear linkage to customer workload readiness. * Define customer-facing abstractions for capacity, node pools, bare metal instances, VM instances, placement policies, OS/runtime images, and reserved compute based on how customers run AI training, inference, and cluster operations. * Partner with engineering and infrastructure operations to define fleet readiness, lifecycle states, failure handling, maintenance workflows, and operational requirements. * Partner with sales, solutions, support, and finance to understand enterprise AI workload requirements and translate them into compute offerings, including reserved capacity, dedicated infrastructure, and VM/bare metal packaging. * Define product requirements for supported compute configurations, including GPU type, node shape, OS/runtime image, network/storage attachment, and compatibility expectations for customer AI workloads. * Improve the reliability, usability, and supportability of CIC compute products across customer onboarding, provisioning, and day-2 operations. * Drive roadmap decisions that improve utilization, reduce stranded capacity, increase enterprise readiness, and lower operational support burden. * Support enterprise customer discovery and roadmap prioritization for bare metal, VMs, reserved capacity, and dedicated infrastructure, with specific attention to customer workload patterns, performance needs, operational workflows, and production readiness. * Create clear product narratives, customer-facing materials, sales enablement, launch plans, and executive updates for enterprise compute offerings. * Track product and business metrics such as usable capacity, allocated capacity, stranded capacity, provisioning time, replacement time, utilization, revenue per deployed GPU, and support burden. Basic Qualifications * 8+ years of product management, technical product management, or equivalent product leadership experience in cloud infrastructure, compute, virtualization, GPU cloud, HPC, private cloud, or enterprise infrastructure platforms. * Experience working backwards from customer workload requirements to define infrastructure products, ideally for AI training, fine-tuning, inference, HPC, data-intensive workloads, or enterprise production systems. * Experience with foundational compute products such as virtual machines, bare metal, cloud instances, node pools, fleet management, capacity management, or infrastructure control planes. * Strong understanding of infrastructure concepts including provisioning, lifecycle management, placement, quota, reservations, OS images, networking, storage attachment, observability, and billing integration. * Familiarity with GPU-based infrastructure and the operational considerations that make compute capacity customer-ready, including drivers, firmware, OS images, high-performance networking, and workload compatibility. * Ability to translate customer AI workload needs, such as distributed training, inference serving, data movement, checkpointing, and cluster operations, into product requirements for compute capacity, lifecycle, observability, and enterprise readiness. * Experience partnering with engineering and infrastructure teams on technically complex systems while driving product outcomes, roadmap decisions, prioritization, and business impact. * Experience supporting enterprise customers, including workload discovery, requirements definition, launch readiness, customer-facing documentation, GTM enablement, and post-launch adoption measurement. * Strong analytical judgment around workload requirements, utilization, capacity planning, product readiness, revenue impact, support cost, and customer adoption. * Strong written and verbal communication skills with engineering, infrastructure operations, finance, sales, support, executive stakeholders, and enterprise customers. Preferred Qualifications * Experience at a hyperscaler, neo-cloud, GPU cloud provider, HPC cloud provider, private cloud platform, or infrastructure SaaS company. * Experience with NVIDIA GPU infrastructure, CUDA, NCCL, OFED, driver compatibility, firmware lifecycle, GPU health/telemetry, or supported GPU software stack management. * Experience with distributed AI training or inference infrastructure, including Slurm, Kubernetes, Ray, model serving platforms, high-performance networking, shared storage, or checkpointing workflows. * Experience with reserved capacity, committed-use contracts, dedicated clusters, savings plans, or enterprise infrastructure commitments. * Experience with bare metal-as-a-service, GPU passthrough virtualization, VM image lifecycle, custom images, or cloud control plane products. Pay & Benefits Our compensation reflects the cost of living across several US geographic markets. At Coupang, your base pay is one part of your total compensation. The base pay for this position ranges from $155,000/year to $267,000/year. Pay is based on several factors including market location and may vary depending on job-related knowledge, skills, and experience. General Description of All Benefits Medical/Dental/Vision/Life, AD&D insurance Flexible Spending Accounts (FSA) & Health Savings Account (HSA) Long-term/Short-term Disability Employee Assistance Program (EAP) program 401K Plan with Company Match 18-21 days of the Paid Time Off (PTO) a year based on the tenure 12 Paid Holidays Up to 6 weeks of Paid Parental leave Pre-tax commuter benefits MTV - [Free] Electric Car Charging Station General Description of Other Compensation “Other Compensation” includes, but is not limited to, bonuses, equity, or other forms of compensation that would be offered to the hired applicant in addition to their established salary range or wage scale. Recruitment Process and Others Recruitment Process * Application Review - Phone Interview - Onsite (or Virtual Onsite) Interview – Offer * The exact nature of the recruitment process may vary according to the specific job and may be changed due to scheduling or other circumstances. * Interview schedules and the results will be informed to the applicant via the e-mail address submitted at the application stage Details to Consider * This job posting may be closed prior to the stated end date for application if all openings are filled. * Coupang has the right to rescind an offer of employment if a candidate is found to have submitted false information as part of the application process. * Those eligible for employment protection (recipients of veteran’s benefits, the disabled, etc.) may receive preferential treatment for employment in accordance with applicable laws. Privacy Notice * Your personal information will be collected and managed by Coupang as stated in the Application Privacy Notice located below: https://www.coupang.jobs/privacy-policy/ Coupang is an equal opportunity employer. All qualified applicants will receive consideration for employment without regard to actual or perceived race (including traits historically associated with race, including but not limited to hair texture and protective hair styles), color, religion, religious creed (including religious dress and grooming practices), sex or gender (including pregnancy, childbirth, breastfeeding, and medical conditions related to pregnancy, childbirth or breastfeeding), gender identity, gender expression, sexual orientation, ,ancestry, national origin (including language use restrictions), age (40 and over), physical or mental disability, medical condition, genetic information, HIV/AIDS or Hepatitis C status, family status (including but not limited to marital or domestic partnership status), military or veteran status, use of a trained dog guide or service animal, political activities or affiliations, ancestry, citizenship, family and medical leave status, status as a victim of any violent crime, or any other characteristic or class protected by the laws or regulations in the locations where we operate. Coupang is also committed to providing a safe work environment for its employees and its consumers. If you need assistance and/or a reasonable accommodation in the application of recruiting process due to a disability, please contact us at usrecruiting@coupang.com. Job Requisition ID: R0059428 Job Requisition ID: R0059428
Isomorphic Labs is applying frontier AI to help unlock deeper scientific insights, faster breakthroughs, and life-changing medicines with an ambition to solve all disease. The future is coming. A future enabled and enriched by the incredible power of machine learning. A future in which diseases are curtailed or cured starting with better and faster drug discovery. Come and be part of an interdisciplinary team driving groundbreaking innovation and play a meaningful role in contributing towards us achieving our ambitious goals, while being a part of an inspiring and collaborative culture. The world we want tomorrow is the one we’re building today. It starts with the culture at this company. It starts with you. ABOUT ISO Isomorphic Labs (IsoLabs) was launched in 2021 to advance human health by building on and beyond the Nobel-winning AlphaFold system. Since then, our interdisciplinary team of drug discovery experts and machine learning specialists has built powerful new predictive and generative AI models that accelerate scientific discovery at digital speed. Our name comes from the belief that there is an underlying symmetry between biology and information science. By harnessing AI’s powerful capabilities, we can use it to model complex biological phenomena to help design novel molecules, anticipate how drugs will perform and develop innovative medicines to treat and cure some of the world’s most devastating diseases. We have built a world-leading drug design engine comprising AI models that are capable of working across multiple therapeutic areas and drug modalities. We are continually innovating on model architecture and developing cutting-edge capabilities to advance rational drug design. Every day, and with each new breakthrough, we’re getting closer to the promise of digital biology, and achieving our ambitious mission to one day solve all disease with the help of AI. YOUR IMPACT We are building the largest foundation models in biotech and applying them immediately to cure disease. You will play a key role and work at a grand scale to deliver the foundations that make this happen. By partnering with in-house machine learning experts and biotech researchers you will join a team to efficiently scale and plan the base on which our groundbreaking AI is built. WHAT YOU WILL DO * You will focus on the end-to-end GPU/TPU (accelerator) strategy, designing infrastructure, optimizing performance, and integrating new hardware to leverage advancements. In partnership with our Machine Learning Platform team, regularly work in the environment to push and support deployments. Regularly be building, monitoring and managing cluster deployments. * Support the technical strategy around hardware acquisition and deployment decisions * Drive research and efficiency design around the infrastructure up to the point of service to the ML platforms teams * Contribute to the efforts for consistently improving the reliability of our ML runs * Operate and handle research, development, and production cloud infrastructure and systems * Partner and collaborate with a diverse set of teams incl. science, research, product, business development and operations * Contribute to core technical decisions (e.g. choice of tooling, infrastructure, and architectural design) SKILLS AND QUALIFICATIONS Essential: * Possess real world experience of large scale AI/ML workloads * Have experience working in cloud compute infrastructure design, preferably GCP * Possess strong programmings skills * Have significant experience working and deploying in Kubernetes * Familiarity with the Nvidia GPU generations Nice to have: * Have a background in either ML SWE or infrastructure SRE work to build on * Have experience leading and delivering projects to multidisciplinary stakeholders * Familiarity with Google TPU generations * Familiarity with: workload scheduling; machine learning efficiency research; familiarity with ML-driven R&D cycles; familiarity with hardware benchmarking CULTURE AND VALUES We are guided by our shared values. It's not about finding people who think and act in the same way. These values help to guide our work and will continue to strengthen it. Thoughtful Thoughtful at Iso is about curiosity, creativity and care. It is about good people doing good, rigorous and future-making science every single day. Brave Brave at Iso is about fearlessness, but it’s also about initiative and integrity. The scale of the challenge demands nothing less. Determined Determined at Iso is the way we pursue our goal. It’s a confidence in our hypothesis, as well as the urgency and agility needed to deliver on it. Because disease won’t wait, so neither should we. Together Together at Iso is about connection, collaboration across fields and catalytic relationships. It’s knowing that transformation is a group project, and remembering that what we’re doing will have a real impact on real people everywhere. CREATING AN EXTRAORDINARY COMPANY We believe that to be successful we need a team with a range of skills and talents. We're building an environment where collaboration is fundamental, learning is shared and every employee feels supported and able to thrive. We value unique experiences, knowledge, backgrounds, and perspectives, and harness these qualities to create extraordinary impact. We are committed to equal employment opportunities regardless of sex, race, religion or belief, ethnic or national origin, disability, age, citizenship, marital, domestic or civil partnership status, sexual orientation, gender identity, pregnancy or related condition (including breastfeeding) or any other basis protected by applicable law. If you have a disability or additional need that requires accommodation, please do not hesitate to let us know. HYBRID WORKING It’s hugely important for us to share knowledge and build strong relationships with each other, and we find it easier to do this if we spend time together in person. This is why we follow a hybrid model, and would require you to be able to come into the office 3 days a week (currently Tuesday, Wednesday, and one other day depending on which team you’re in). If you have additional needs that would prevent you from following this hybrid approach, we’d be happy to talk through these if you’re selected for an initial screening call. Please note that when you submit an application, your data will be processed in line with our privacy policy. >> Click to view other open roles at Isomorphic Labs