
Coupang · Seattle
Company Introduction We exist to wow our customers. We know we’re doing the right thing when we hear our customers say, “How did we ever live without Coupang...
Company Introduction
We exist to wow our customers. We know we’re doing the right thing when we hear our customers say, “How did we ever live without
Coupang?” Born out of an obsession to make shopping, eating, and living easier than ever, we’re collectively disrupting the
multi-billion-dollar e-commerce industry from the ground up. We are one of the fastest-growing e-commerce companies that
established an unparalleled reputation for being a dominant and reliable force in South Korean commerce.
We are proud to have the best of both worlds — a startup culture with the resources of a large global public company. This fuels
us to continue our growth and launch new services at the speed we have been since our inception. We are all entrepreneurs
surrounded by opportunities to drive new initiatives and innovations. At our core, we are bold and ambitious people that like to
get our hands dirty and make a hands-on impact. At Coupang, you will see yourself, your colleagues, your team, and the company
grow every day.
Our mission to build the future of commerce is real. We push the boundaries of what’s possible to solve problems and break
traditional tradeoffs. Join Coupang now to create an epic experience in this always-on, high-tech, and hyper-connected world.
About Us
Coupang is at the forefront of the AI and high-performance computing (HPC) revolution. We are building a next-generation cloud
platform designed to provide developers, researchers, and enterprises with seamless, scalable, and powerful access to accelerated
computing. As the demand for AI, machine learning, and data-intensive workloads skyrockets, we are looking for a visionary product
leader to define the future of our core compute offerings.
Job Overview
CIC is looking for a Group Product Manager to own the foundational compute platform that powers enterprise AI workloads. This role
spans fleet management, capacity management, bare metal, virtualized compute, node lifecycle, placement, reservations, and
infrastructure-level customer experience.
The Product Manager will be responsible for turning physical GPU and CPU capacity into reliable, customer-ready, observable, and
billable compute products. This includes defining how capacity is reserved, provisioned, validated, monitored, maintained,
packaged, and exposed to enterprise customers.
This role sits at the compute foundation layer. It enables higher-level orchestration and workload services such as Kubernetes,
Slurm, Ray, jobs, notebooks, and inference endpoints. The candidate should understand how those orchestration and workload systems
depend on foundational compute infrastructure.
Key Responsibilities
working backwards from customer AI workload requirements.
lifecycle actions, observability, billing integration, maintenance, and deprecation, with clear linkage to customer workload
readiness.
OS/runtime images, and reserved compute based on how customers run AI training, inference, and cluster operations.
maintenance workflows, and operational requirements.
compute offerings, including reserved capacity, dedicated infrastructure, and VM/bare metal packaging.
network/storage attachment, and compatibility expectations for customer AI workloads.
day-2 operations.
operational support burden.
infrastructure, with specific attention to customer workload patterns, performance needs, operational workflows, and production
readiness.
enterprise compute offerings.
replacement time, utilization, revenue per deployed GPU, and support burden.
Basic Qualifications
infrastructure, compute, virtualization, GPU cloud, HPC, private cloud, or enterprise infrastructure platforms.
fine-tuning, inference, HPC, data-intensive workloads, or enterprise production systems.
management, capacity management, or infrastructure control planes.
OS images, networking, storage attachment, observability, and billing integration.
including drivers, firmware, OS images, high-performance networking, and workload compatibility.
and cluster operations, into product requirements for compute capacity, lifecycle, observability, and enterprise readiness.
roadmap decisions, prioritization, and business impact.
customer-facing documentation, GTM enablement, and post-launch adoption measurement.
support cost, and customer adoption.
stakeholders, and enterprise customers.
Preferred Qualifications
company.
supported GPU software stack management.
high-performance networking, shared storage, or checkpointing workflows.
commitments.
plane products.
Pay & Benefits
Our compensation reflects the cost of living across several US geographic markets. At Coupang, your base pay is one part of your
total compensation.
The base pay for this position ranges from $155,000/year to $267,000/year. Pay is based on several factors including market
location and may vary depending on job-related knowledge, skills, and experience.
General Description of All Benefits
Medical/Dental/Vision/Life, AD&D insurance
Flexible Spending Accounts (FSA) & Health Savings Account (HSA)
Long-term/Short-term Disability
Employee Assistance Program (EAP) program
401K Plan with Company Match
18-21 days of the Paid Time Off (PTO) a year based on the tenure
12 Paid Holidays
Up to 6 weeks of Paid Parental leave
Pre-tax commuter benefits
MTV - [Free] Electric Car Charging Station
General Description of Other Compensation
“Other Compensation” includes, but is not limited to, bonuses, equity, or other forms of compensation that would be offered to the
hired applicant in addition to their established salary range or wage scale.
Recruitment Process and Others
Recruitment Process
other circumstances.
Details to Consider
the application process.
treatment for employment in accordance with applicable laws.
Privacy Notice
Coupang is an equal opportunity employer. All qualified applicants will receive consideration for employment without regard to
actual or perceived race (including traits historically associated with race, including but not limited to hair texture and
protective hair styles), color, religion, religious creed (including religious dress and grooming practices), sex or gender
(including pregnancy, childbirth, breastfeeding, and medical conditions related to pregnancy, childbirth or breastfeeding), gender
identity, gender expression, sexual orientation, ,ancestry, national origin (including language use restrictions), age (40 and
over), physical or mental disability, medical condition, genetic information, HIV/AIDS or Hepatitis C status, family status
(including but not limited to marital or domestic partnership status), military or veteran status, use of a trained dog guide or
service animal, political activities or affiliations, ancestry, citizenship, family and medical leave status, status as a victim
of any violent crime, or any other characteristic or class protected by the laws or regulations in the locations where we operate.
Coupang is also committed to providing a safe work environment for its employees and its consumers. If you need assistance and/or
a reasonable accommodation in the application of recruiting process due to a disability, please contact us
at usrecruiting@coupang.com.
Job Requisition ID: R0059428
Job Requisition ID: R0059428
Ready to do the most impactful work of your career? At Coinbase, we are uncompromising on our mission to increase economic freedom. The bar is high, the environment is intense, and we like it that way. This isn't a place for complacency, it’s a place to be pushed past your perceived limits. If you're ready to build the future of finance alongside people who refuse to settle for "good enough," you belong here. Coinbase is a remote-first, but not remote-only company. Expect to get together quarterly for intense in-person working sessions called “surges.” learn more about working at Coinbase. The Core Infrastructure team within Coinbase's Platform product group builds the foundational systems that keep Coinbase online, secure, and scalable, owning the compute and networking platforms that power every product and service across the company. As the Group Product Manager for Core Infrastructure & Reliability, you'll own the product vision and multi-year strategy for Coinbase's cloud infrastructure, driving the design, operation, and scaling of the systems that underpin hundreds of billions of dollars in annual transaction volume. You'll partner deeply with Engineering, SRE, Security, and Finance to ensure Coinbase's infrastructure is reliable, cost-efficient, and resilient across multiple cloud environments and regions. What you’ll do: * Own the product strategy and roadmap for Core Infrastructure, spanning compute, networking, multi-region and multi-cloud architecture, and platform reliability. * Strengthen infrastructure reliability and resilience programs, defining platform-level SLOs, capacity planning, failover capabilities, and incident reduction targets to meet the uptime demands of a global financial platform. * Lead evaluation and adoption of cloud infrastructure technologies (Kubernetes, service mesh, distributed storage, observability, infrastructure-as-code), making build-vs-buy decisions that balance cost, speed, and long-term scalability. * Align infrastructure investments with business priorities, optimize cloud spend, and ensure infrastructure services meet regulatory and compliance requirements across operating jurisdictions. * Shape how Coinbase's product and engineering teams consume infrastructure by building self-serve capabilities, improving developer experience, and reducing friction across provisioning, deployment, and observability workflows. * Drive infrastructure cost strategy by analyzing utilization patterns, identifying the highest-leverage optimization opportunities, and translating infrastructure metrics into executive-level investment cases. Required Skills and Experience: * 10+ years of product management experience, with 5+ years focused on cloud infrastructure, platform engineering, or infrastructure reliability at scale. * Track record of defining and executing multi-year infrastructure platform strategies that delivered measurable improvements in reliability, cost efficiency, or engineering velocity. * Technical depth to engage with engineers on compute, networking, storage, Kubernetes, multi-cloud architecture, and distributed systems design. * Demonstrated ability to lead complex, cross-functional initiatives in regulated, high-trust environments. * Proven ability to articulate complex infrastructure strategy, cost tradeoffs, and long-term platform bets to senior and executive leadership. * Utilizes generative AI responsibly, maintaining human oversight to deliver business-ready outputs and drive measurable improvements in workflow efficiency, cost, and quality. Pay Transparency Notice: Base salary varies by location (see range below). Total compensation may also include equity and bonus eligibility, and benefits (medical, dental, vision, 401(k)). Annual base salary range (excluding equity and bonus): $243,865—$286,900 USD * Application Limit: Candidates may submit a maximum of 3 applications within a 6-month period. * Equal Opportunity Employer: Coinbase is an Equal Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, disability, protected veteran status, or genetic information. Applicants with criminal histories will be considered consistent with applicable federal, state, and local laws. * US Applicants: View Employee Rights, Know Your Rights, and E-Verify Notice of Participation. * Accommodations: If you are an individual with a disability who needs a reasonable accommodation, email us your request and contact info at accommodations[at]coinbase.com. Need screen reading technology? Click here to download a free compatible screen reader and view the tutorial. * Data Privacy & Arbitration: By submitting your application, you agree to our Candidate Privacy Notice. US applicants: By submitting your application, you agree to Arbitration of Disputes.
WHO WE ARE ABOUT STRIPE Stripe is a financial infrastructure platform for businesses. Millions of companies — from the world's largest enterprises to the most ambitious startups — use Stripe to accept payments, grow their revenue, and accelerate new business opportunities. Our mission is to increase the GDP of the internet, and we have a staggering amount of work ahead. That means you have an unprecedented opportunity to put the global economy within everyone's reach while doing the most important work of your career. ABOUT THE TEAM The Core Change Management group is responsible for the systems that let every Stripe engineer ship code, configuration, and infrastructure changes safely and at high velocity. You will be embedded primarily on the Service Deployments team — the owners of Stripe's end-to-end code deployment platform — with regular collaboration with the Resource Automation and Feature Deployments teams. Service Deployments owns the full lifecycle of software changes at Stripe. The team's mission is to let developers roll out code and configuration changes safely without sacrificing productivity, with a goal of meaningfully reducing change-related production incidents year over year. The team operates a meaningful on-call rotation and owns the systems that sit in the critical path of every engineer's daily workflow at Stripe. Resource Automation owns the safe-by-default infrastructure change layer: automated remote execution, incremental Infrastructure as Code tooling, cloud resource inventory, and cloud account governance and IAM role management. You will collaborate with this team on projects that span the boundary between deployment orchestration and cloud resource management. Feature Deployments owns Stripe's feature flag system, merchant entitlements, configuration management and distribution, and the audit log of change-correlated events. You will work with this team when deployment pipelines intersect with feature rollout and change safety tooling. WHAT MAKES THIS ROLE COMPELLING * You own the foundation of how Stripe ships software. The deployment platform sits in the critical path of every engineer's workflow at Stripe. The decisions you make affect thousands of deploys per day across hundreds of services, directly determining how fast and safely Stripe's product evolves. * Technically rich, architecturally active. The team is executing several concurrent platform transformations: containerizing host-based services at scale, adding intelligent multi-service deploy pipelines, extending real-time anomaly detection to earlier stages of traffic shifts, and rebuilding deployment event infrastructure on top of a durable message bus. This is not maintenance work — the architecture is in motion. * Broad surface area, real ownership. You will span the full stack from container scheduling and deployment orchestration business logic to the developer-facing internal platform UI. The problems are multi-layered: reliability, developer experience, performance, and safety all at once. * Your judgment prevents incidents. The team's explicit goal is to drive down change-related incidents across Stripe by building better detection, smarter pipelines, and safer defaults. Your technical decisions have a direct and measurable safety impact on Stripe's reliability. * Agency to shape technical strategy. As a Staff engineer on Service Deployments, you will set technical direction for the team's systems, author designs that span multiple teams, and be the person engineering managers and engineers turn to for the hardest deployment infrastructure questions. RESPONSIBILITIES * Own end-to-end technical delivery of large, ambiguous infrastructure projects — from initial design through production launch and long-term reliability. Author the design, sequence the work, unblock the team, and shepherd projects to landed impact. * Architect the next generation of Stripe's deployment platform. Lead technical design of the deployment orchestrator's evolution — including multi-service dependency-aware autodeploy pipelines, Kubernetes-native deployment primitives, and fleetwide container migration — defining the API contracts, rollout strategies, and operational model that hundreds of teams depend on. * Extend deploy anomaly detection. Evolve blue-green traffic analysis: extend coverage to earlier traffic-split stages, design API/method-based regression detection, and build a self-service onboarding system that makes anomaly detection the default for all supported service types. * Lead the host-to-container fleet migration. Drive sequencing, backward compatibility, and cross-team coordination for migrating Stripe's fleet of host-based services to containerized, fleetwide deployments — keeping the production deployment system operational while executing the transformation. * Own reliability and operational excellence for the deployment platform. Lead incident response; systematically reduce operational toil; and make reliability, security, and maintainability first-class properties of the systems you own. * Build deployment event infrastructure. Own the deployment notification and event-publishing architecture — designing the event schema, durability model, and integration contracts that downstream systems rely on for observability and automation. * Collaborate across Core Change Management. Partner with Resource Automation on projects that span deployment orchestration and cloud resource management (IAM, account provisioning, infrastructure automation), with Feature Deployments on change-safety tooling (feature flags, configuration management, change audit logs) that integrates with or depends on the deployment pipeline, and with the service mesh team on routing capabilities that enable advanced deployment patterns such as canary rollouts and merchant-priority traffic shaping. * Set the technical bar. Own critical design reviews, establish standards for deployment safety and developer experience, mentor senior engineers through high-stakes architectural decisions, and advocate for the right abstractions — code that consuming teams can adopt without becoming deployment infrastructure experts. * Decompose complexity for the team. Translate large, open-ended platform challenges into scoped, parallelizable work; help engineers grow by framing problems clearly and providing decisive technical guidance on the hardest questions. WHO YOU ARE MINIMUM REQUIREMENTS * 10+ years of professional software engineering experience, with a demonstrated track record of designing and shipping production infrastructure systems of significant scale and complexity. * Proven ability to lead large, ambiguous infrastructure projects end-to-end — from technical design through delivery — including managing cross-team dependencies and coordinating migrations across many consuming teams. * Deep expertise in distributed systems and deployment orchestration: strong foundations in how services are built, scheduled, and operated at scale, including rollout strategies, staged delivery, and failure modes. * Hands-on experience with Kubernetes and container-based deployments, including service lifecycle management, workload scheduling, and the operational challenges of migrating large fleets from VM-based to containerized infrastructure. * Strong background in service reliability and operational excellence: demonstrated ability to lead incident response, reduce toil, and build systems that are reliable, debuggable, and maintainable by a team. * Track record of broad technical impact across multiple large systems: fluency across a complex codebase, force-multiplier effect through code review and mentorship, and the ability to set technical direction for a team rather than just execute within it. PREFERRED REQUIREMENTS * Background in deployment safety systems: anomaly detection, automated rollback, progressive delivery, or similar mechanisms that reduce the blast radius of bad deployments. * Familiarity with event-driven architectures (Kafka or equivalent) applied to deployment lifecycle observability and notification. * Experience with Infrastructure as Code at scale — Terraform or equivalent — particularly in the context of cloud resource governance and IAM management in AWS or Azure. * Developer platform or internal tooling background: a strong developer experience sensibility and the ability to build abstractions that reduce toil for the engineering teams that depend on your platform. * Change management and feature rollout systems: experience with feature flags, configuration distribution, or audit-log infrastructure that provides safety guardrails around production changes. * Familiarity with service mesh concepts (canary deployments, weighted routing, traffic-splitting) sufficient to collaborate effectively with partner teams on routing capabilities that enable advanced deployment patterns. IN-OFFICE EXPECTATIONS Office-assigned Stripes in most of our locations are currently expected to spend at least 50% of the time in a given month in their local office or with users. This expectation may vary depending on role, team and location. For example, Stripes in Stripe Delivery Center roles in Mexico City, Mexico, Bengaluru, India, and Dublin, Ireland work 100% from the office. Also, some teams have greater in-office attendance requirements, to appropriately support our users and workflows, which the hiring manager will discuss. This approach helps strike a balance between bringing people together for in-person collaboration and learning from each other, while supporting flexibility when possible.
ABOUT US Graphcore is one of the world’s leading innovators in Artificial Intelligence compute. It is developing hardware, software and systems infrastructure that will unlock the next generation of AI breakthroughs and power the widespread adoption of AI solutions across every industry. As part of the SoftBank Group, Graphcore is a member of an elite family of companies responsible for some of the world’s most transformative technologies. Together, they share a bold vision: to enable Artificial Super Intelligence and ensure its benefits are accessible to everyone. Graphcore’s teams are drawn from diverse backgrounds and bring a broad range of skills and perspectives. A melting pot of AI research specialists, silicon designers, software engineers and systems architects, Graphcore enjoys a culture of continuous learning and constant innovation. JOB SUMMARY We are looking for a Staff Digital Twin Engineer to help design, build and evolve digital twin capabilities that support the development, validation and operation of advanced AI compute systems. Reporting to an Engineering Manager within the Digital Twin / Simulation Platform team, this role will lead significant technical work across simulation, 3D visualization, data integration and workflow automation, with a particular focus on the Omniverse platform, OpenUSD-based pipelines, data center operations, and rack and blade-level design and simulation workflows. THE TEAM The Digital Twin / Simulation Engineer works with internal Graphcore teams and Softbank Group affiliates. Team will be responsible for the design, development and integration of simulation and visualization capabilities that support data center digital twin, machine behavior modeling and interactive 3D experiences. This role combines simulation engineering, real-time visualization and 3D scene-building using NVIDIA Omniverse and modern tool chains. Team will build behavioral logic, physics integrations, and visualization layers that enable accurate, high-performance representation of data center components, job sites and workflows. RESPONSIBILITIES AND DUTIES * This is a hands-on technical leadership role, involving architecture, implementation, technical guidance and mentoring, without direct line management responsibility. * Apply MBSE methodologies to drive the development of systems and ensure traceability from requirements to implementation. * Develop and deploy simulation logic and behavioral models governing machine and system behavior. Implement real‑time simulation algorithms, including kinematics, dynamics, and control‑system interactions. This includes integration of physics engines (e.g., PhysX, Omniverse Physics, or equivalent) to support realistic behavior and machine–environment interactions. * Develop and execute model validation and verification plans, ensuring digital twins are accurate and numerically stable. * Develop robust scene composition, data modeling and asset management approaches for engineering digital twins across component, blade, rack and facility-level views. * Create high-fidelity 3D models, scenes and environments in Omniverse. Develop interfaces that connect simulations to real‑time visualization, dashboards, and operator‑focused experiences. Implement automated test pipelines and document model correlations, limitations, and validation scoring. * Contribute to toolchain and pipeline design, including scripting utilities, automated test frameworks, and modular code libraries. Collaborate across data engineering/architecture, autonomy, product groups, and engineering to ensure interoperability and alignment with business needs. * Evaluate and implement new simulation libraries and real‑time rendering technologies that enhance simulation fidelity and performance. Work hands‑on with multiple digital‑twin platforms, such as: NVIDIA Omniverse (primary environment), CARLA, Siemens NX/Teamcenter/Plant Simulation, Dassault 3DEXPERIENCE, Ansys Twin Builder, Unity or similar. * Integrate engineering and operational data from simulation, test, telemetry, design systems, asset inventories and infrastructure sources into coherent digital twin environments. * Support testing, debugging, performance analysis and reliability improvements across digital twin applications and services. * Establish engineering best practices for OpenUSD schemas, Omniverse workflows, version control, documentation and release processes. * Provide technical guidance, code review and mentoring to engineers working with digital twin, simulation and data integration technologies. * Produce clear technical documentation and communicate progress, risks and decisions to technical and non-technical stakeholders. CANDIDATE PROFILE ESSENTIAL * Degree in Mechanical, Electrical, Systems Engineering, Computer Science or related discipline. * 8+ years of engineering experience in automation, simulation, or controls. * Experience designing, building or maintaining digital twin, simulation, 3D visualization or engineering platform software. * Hands-on experience with the Omniverse platform, including Omniverse Kit, extensions, connectors or related development workflows. * Strong knowledge of OpenUSD / USD concepts, including composition, layering, schemas, assets and scene graph workflows. * Experience working with data center operations, infrastructure environments, operational planning or systems used to manage complex technical facilities. * Experience developing simulation or modeling workflows that support engineering decisions for complex hardware, infrastructure or operational systems. * Strong software engineering skills in Python, C++ or both, with experience building maintainable production-quality tools or services. * Experience integrating complex engineering and operational data from multiple systems, formats or APIs into usable software workflows. * Understanding simulation, real-time visualization, physics-based modeling, rendering pipelines or spatial computing concepts. * Ability to lead complex technical tasks independently, manage priorities and make sound engineering decisions with incomplete information. * Strong collaboration and communication skills, with the ability to work effectively across multiple engineering disciplines. * Experience reviewing designs and code, sharing technical knowledge and helping other engineers adopt new tools and practices. DESIRABLE * Experience applying digital twin technologies in data center, high-performance computing, semiconductor, robotics, manufacturing or complex systems environments. * Experience with rack and blade design concepts, including physical layout, serviceability, power, cooling, cabling, networking or mechanical constraints. * Experience modeling thermal, power, airflow, mechanical, cabling, networking or serviceability characteristics for rack-scale or facility-scale systems. * Experience with synthetic data generation, sensor simulation, telemetry visualization or operational monitoring workflows. * Knowledge of containerized deployment, cloud or on-premises infrastructure, CI/CD and automated testing for engineering tools. * Experience with 3D asset pipelines, materials, lighting, rendering, physics simulation or model optimization. * Familiarity with adjacent platforms or technologies such as Unreal Engine, Unity, Blender, ROS, CAD/CAE tools, DCIM tools or PLM systems. * Understanding of AI, machine learning, accelerated computing or large-scale system architecture concepts. * Experience developing user-facing tools for engineering or operations teams, including workflow design, usability improvements and technical enablement. BENEFITS In addition to a competitive salary, Graphcore offers flexible working and a comprehensive benefits package designed to support your health, wellbeing and financial future. Our benefits include medical, dental and vision coverage, Flexible Spending Accounts (FSAs), Health Savings Accounts (HSAs), disability and life insurance, a 401(k) retirement plan, commuter benefits, wellness services and an Employee Assistance Programme (EAP). We welcome people of different backgrounds and experiences; we're committed to building an inclusive work environment that makes Graphcore a great home for everyone. We offer an equal opportunity process and understand that there are visible and invisible differences in all of us. We can provide a flexible approach to interview and encourage you to chat to us if you require any reasonable adjustments.