
Coupang · Seoul
About Coupang We exist to wow our customers. We know we’re doing the right thing when we hear our customers say, “How did we ever live without Coupang?” Born o...
About Coupang
We exist to wow our customers. We know we’re doing the right thing when we hear our customers say, “How did we ever live without
Coupang?” Born out of an obsession to make shopping, eating, and living easier than ever, we’re collectively disrupting the
multi-billion-dollar e-commerce industry from the ground up. We are one of the fastest-growing e-commerce companies that
established an unparalleled reputation for being a dominant and reliable force in South Korean commerce.
We are proud to have the best of both worlds — a startup culture with the resources of a large global public company. This fuels
us to continue our growth and launch new services at the speed we have been since our inception. We are all entrepreneurial
surrounded by opportunities to drive new initiatives and innovations. At our core, we are bold and ambitious people that like to
get our hands dirty and make a hands-on impact. At Coupang, you will see yourself, your colleagues, your team, and the company
grow every day.
Our mission to build the future of commerce is real. We push the boundaries of what’s possible to solve problems and break
traditional tradeoffs. Join Coupang now to create an epic experience in this always-on, high-tech, and hyper-connected world.
Role Overview
We are seeking a Sr. Staff Observability Engineer to lead the design and evolution of our observability platform for
a GPU-as-a-Service (GPUaaS) infrastructure. This role will own the end-to-end telemetry strategy—from high-throughput metric
ingestion to log pipelines and real-time visualization—powering deep insights into GPU clusters, datacenter systems, and
distributed workloads.
You will architect and operate planet-scale telemetry pipelines leveraging Grafana Alloy, Mimir, Loki, and Vector, ensuring
high-fidelity observability across GPU workloads, Kubernetes clusters, and datacenter infrastructure.
Key Responsibilities
to SLO-driven, predictive, and automated observability.
Qualifications & Requirements
Core Impact
Recruitment Process and Others
Recruitment Process
other circumstances.
stage.
Details to Consider
the application process.
treatment for employment in accordance with applicable laws.
communicated to the candidate at the appropriate time before the offer.
be either skipped, shortened or extended if necessary for business purposes.
Privacy Notice
below.
Document Return Policy
1. This notification is given pursuant to Article 11 (6) of the Fair Hiring Procedure Act.
2. A job applicant, who has applied but not been finally selected for a position at Coupang (the “Company”), may request the
Company to return his/her hiring documents submitted pursuant to the Fair Hiring Procedure Act. However, this will not apply
where the hiring documents were submitted via the website of the Company or e-mail, or where the job applicant submitted those
documents voluntarily without a request from the Company. In addition, if the hiring documents were destroyed due to a natural
disaster or any other reasons not attributable to the Company, such documents will be deemed to have been returned to the job
applicant.
3. A job applicant who wishes to request the return of his/her hiring documents pursuant to the main sentence of paragraph 2
above should fill out a “Request for Return of Hiring Documents” [Annex Form No. 3 in the Enforcement Rule of the Fair Hiring
Procedure Act] and submit It by email (recruitingops@coupang.com). In such case, within fourteen (14) days from the date of
identifying the receipt of the request, the Company will send the hiring documents to the job applicant’s designated address
via registered mail. Please be informed that the job applicant is required to pay the postage on the registered mail.
4. In preparation for a job applicant’s request for the return of hiring documents pursuant to the main sentence of paragraph 2
above, the Company shall retain the original hiring documents submitted by the job applicant for 180 days from the completion
of the recruiting process. If no request is made until the end of this period, all his/her hiring documents will be destroyed
immediately in accordance with the Personal Information Protection Act.
5. The above paragraphs 1 - 4 shall only apply when the labor-related laws of Korea govern the application. They are otherwise
not applicable.
Company Introduction We exist to wow our customers. We know we’re doing the right thing when we hear our customers say, “How did I ever live without Coupang?” Born out of an obsession to make shopping, eating, and living easier than ever, we’re collectively disrupting the multi-billion-dollar e-commerce industry from the ground up. We are one of the fastest-growing e-commerce companies that established an unparalleled reputation for being a dominant and reliable force in South Korean commerce. We are proud to have the best of both worlds — a startup culture with the resources of a large global public company. This fuels us to continue our growth and launch new services at the speed we have been since our inception. We are all entrepreneurs surrounded by opportunities to drive new initiatives and innovations. At our core, we are bold and ambitious people that like to get our hands dirty and make a hands-on impact. At Coupang, you will see yourself, your colleagues, your team, and the company grow every day. Our mission to build the future of commerce is real. We push the boundaries of what’s possible to solve problems and break traditional trade-offs. Join Coupang now to create an epic experience in this always-on, high-tech, and hyper-connected world. Role Overview We are seeking a Sr Staff System Engineer, GPU Fleet for our Coupang Intelligent Cloud (CIC) team, to serve as the senior technical owner for our hyperscale GPU compute infrastructure. In this role, you will define fleet architecture, drive reliability and automation at scale, and lead the operation and evolution of GPU systems supporting large‑scale AI training and inference workloads. This is a hands‑on, staff‑level individual contributor role with broad technical ownership, high operational impact, and significant cross‑functional influence across hardware, infrastructure, and datacenter operations. CIC builds the infrastructure for abundant intelligence. We partner with leading AI labs, governments, and enterprises to deliver hyperscale GPU compute with high reliability, performance, and efficiency. Our infrastructure supports some of the most demanding AI training and inference workloads in production today. We operate with urgency, deep ownership, and a strong bias toward execution. Reliability, operational excellence, and rigorous systems engineering are core to our business. What You Will Do As a Sr Staff System Engineer, GPU Fleet, you will be the senior technical owner for CIC’s large‑scale GPU compute infrastructure. This is a hands‑on senior individual contributor role with fleet‑level responsibility and broad cross‑functional influence. You will define the technical direction for how GPU fleets are architected, operated, automated, and evolved across multiple generations of hardware. Your work will directly affect fleet reliability, operating efficiency, scalability, and customer success. This role does not involve people management, but it carries principal‑level scope, autonomy, and decision‑making authority across infrastructure, hardware, and operations. Key Responsibilities: Fleet Architecture & Technical Ownership * Own the end‑to‑end technical architecture of hyperscale GPU fleets, including hardware platform selection, firmware strategy, OS configuration, drivers, networking, and observability. * Define and enforce technical standards and best practices for fleet reliability, availability, performance, and operability. * Lead major fleet‑wide initiatives such as new GPU platform bring‑ups, multi‑generation hardware transitions, and architectural redesigns. * Evaluate trade‑offs across cost, performance, reliability, and time‑to‑deploy, and make technically sound decisions under ambiguity. Reliability, Availability & Performance * Set and drive fleet‑level reliability, availability, and performance objectives. * Lead root‑cause analysis and resolution of complex, systemic failures affecting large portions of the fleet or multiple datacenters. * Identify recurring failure patterns and drive long‑term fixes spanning hardware, software, automation, and operational processes. * Work directly with hardware vendors and partners to resolve platform‑level issues and influence future hardware designs. Automation & Systems Engineering * Design and build large‑scale automation systems for: * GPU fleet provisioning and lifecycle management * GPU health validation, diagnostics, and certification * Automated remediation, recovery, and replacement workflows * Eliminate manual operational toil through durable, well‑designed tooling that scales with fleet growth. * Ensure all fleet systems are observable, testable, and resilient under failure conditions. Operational Leadership * Act as a senior escalation point for critical production incidents impacting GPU availability or customer workloads. * Participate in on‑call rotations with a strong emphasis on preventing future incidents, not just responding to them. * Lead high‑severity post‑incident reviews and ensure learnings are translated into concrete engineering and process improvements. Technical Influence & Mentorship * Provide technical mentorship and guidance to system and infrastructure engineers across the organization. * Serve as a trusted technical partner to platform engineering, networking, datacenter operations, and leadership teams. * Influence CIC’s long‑term infrastructure roadmap through strong technical judgment and data‑driven recommendations. Basic Qualifications * 12+ Years of overall experience with at least 8+ years of experience in Linux systems engineering, infrastructure engineering, or datacenter operations, operating production environments with strict uptime and performance requirements. * Deep, hands‑on expertise in Linux system internals, including process scheduling, memory management, filesystem behavior, networking, kernel behavior, and system performance analysis. * Demonstrated experience operating hardware‑intensive infrastructure in production, including bare‑metal servers at scale. * Proven ability to debug complex issues across multiple system layers, including hardware components, firmware/BIOS, kernel drivers, OS configuration, and user‑space services. * Extensive experience writing production‑grade automation using Python and Bash for provisioning, configuration management, diagnostics, remediation, and fleet operations. * Strong understanding of how to design systems that are observable, resilient, and safe under failure, rather than reliant on manual intervention. Preferred Qualifications * Direct experience operating large‑scale GPU fleets supporting AI/ML training and/or inference workloads in production. * Familiarity with modern GPU platforms and ecosystems, including GPU drivers, CUDA, NCCL, and high‑performance compute workloads. * Experience with high‑speed interconnects and datacenter networking, such as NVLink, InfiniBand, RDMA, and high‑throughput Ethernet. * Prior ownership of fleet‑wide or platform‑wide initiatives, such as new hardware bring‑ups, major architectural changes, or reliability transformations. * Experience partnering directly with hardware vendors or manufacturers to troubleshoot systemic issues or influence future platform designs. * Strong intuition for failure modes at scale, including cascading failures, correlated faults, and second‑order effects across systems. * History of acting as a technical authority or escalation point for ambiguous, high‑impact production problems. * Ability to mentor engineers through design reviews, technical problem solving, and modelling strong operational ownership. * Experience participating in on‑call rotations and responding to high‑severity production incidents with clear ownership, urgency, and technical leadership. * Strong written and verbal communication skills, including clear post‑incident reviews and technical documentation. Type of work model Hybrid Details to consider * Those eligible for employment protection (recipients of veteran’s benefits, the disabled, etc.) may receive preferential treatment for employment in accordance with applicable laws. Privacy Notice * Your personal information will be collected and managed by Coupang as stated in the Application Privacy Notice located below. https://privacy.coupang.com/en/land/jobs/
About Pinterest: Millions of people around the world come to our platform to find creative ideas, dream about new possibilities and plan for memories that will last a lifetime. At Pinterest, we’re on a mission to bring everyone the inspiration to create a life they love, and that starts with the people behind the product. Discover a career where you ignite innovation for millions, transform passion into growth opportunities, celebrate each other’s unique experiences and embrace the flexibility to do your best work. Creating a career you love? It’s Possible. At Pinterest, AI isn't just a feature, it's a powerful partner that augments our creativity and amplifies our impact, and we’re looking for candidates who are excited to be a part of that. To get a complete picture of your experience and abilities, we’ll explore your foundational skills and how you collaborate with AI. Through our interview process, what matters most is that you can always explain your approach, showing us not just what you know, but how you think. You can read more about our AI interview philosophy and how we use AI in our recruiting process here. Pinterest is looking for a Senior Staff Software Engineer to lead Performance Ads Formats, a team building ad products and experiences that help Pinners turn intent into action and help lower-funnel advertisers drive stronger outcomes. In this role, you’ll be a senior technical leader across a suite of products such as Deal Ads, dynamic ad experiences, post ad click journeys, and more. You’ll partner closely with Product, Design, Data Science, PMM, TPM, and Engineering leaders to turn ambiguous opportunities into scalable, reliable ad products with measurable business impact. What you’ll do: * Lead technical strategy and architecture across Performance Ads Formats, with a focus on growing good clicks, improving click quality, increasing conversions, and enhancing the Pinner experience. * Drive cross-functional product initiatives from discovery through launch and readout, including workstreams such as post-click ad journeys, dynamic ad formats, deal ads, etc. * Design scalable, resilient, and maintainable systems across client, API, backend, and partner surfaces, biasing for impact and balancing speed, quality, privacy / compliance needs, and long-term ownership. * Partner with Product and Data Science to make data-driven prioritization and launch decisions, using experiment results, guardrails, and business impact to decide what to scale, iterate, or stop. * Raise the production quality bar through strong design reviews, code reviews, testing strategy, QA / bug bash processes, observability, incident readiness, and thoughtful tech debt reduction. * Mentor senior engineers and create technical strategy docs, design docs, code, and analysis artifacts that become examples of clarity, simplicity, and quality across multiple teams. * Use pragmatic tools, including AI where useful, to accelerate knowledge discovery, prototyping, documentation, and quality checks while applying strong judgment and verification. What we’re looking for: * Deep technical architecture expertise in high-scale product systems, with the ability to reason from first principles and dig deep into how complex technical systems work under the hood. * A track record of leading technically complex, ambiguous, multi-team initiatives that shipped measurable product or business impact. * Excellent cross-functional collaboration and communication skills, including the ability to align senior stakeholders, make tradeoffs explicit, and influence without authority. * Strong data-driven decision making and prioritization: you can explain why something is the right bet, what evidence supports it, and how you would know whether it worked. * A production-quality mindset for scalable, reliable, maintainable software, including testing, observability, incident response, operational cost, and long-term system health. * Experience mentoring Senior and Staff engineers and raising the technical bar through reviews, architecture guidance, knowledge sharing, and crisp technical writing. * Demonstrated ability to use AI to improve speed and quality in your day-to-day workflow for relevant outputs. * Strong track record of critical evaluation and verification of AI-assisted work (e.g., testing, source-checking, data validation, peer review).High integrity and ownership. * you protect sensitive data, avoid over-reliance on AI, and remain accountable for final decisions and deliverables. * Nice to have: experience in ads, e-commerce, recommendations, experimentation platforms, or practical AI-assisted engineering workflows. * Bachelor’s/Master’s degree in a relevant field such as Computer Science, or 8+ YOE as a Software Engineer. Relocation Statement: * This position is not eligible for relocation assistance. Visit our PinFlex page to learn more about our working model. In-Office Requirement Statement: * We let the type of work you do guide the collaboration style. That means we're not always working in an office, but we continue to gather for key moments of collaboration and connection. * This role will need to be in the office for in-person collaboration once a week and therefore needs to be in a commutable distance from one of the following offices: San Francisco or Palo Alto offices. #LI-HYBRID #LI-KBF At Pinterest we believe the workplace should be equitable, inclusive, and inspiring for every employee. In an effort to provide greater transparency, we are sharing the base salary range for this position. The position is also eligible for equity. Final salary is based on a number of factors including location, travel, relevant prior experience, or particular skills and expertise. Information regarding the culture at Pinterest and benefits available for this position can be found here. US based applicants only $245,402—$429,454 USD Our Commitment to Inclusion: Pinterest is an equal opportunity employer and makes employment decisions on the basis of merit. We want to have the best qualified people in every job. All qualified applicants will receive consideration for employment without regard to race, color, ancestry, national origin, religion or religious creed, sex (including pregnancy, childbirth, or related medical conditions), sexual orientation, gender, gender identity, gender expression, age, marital status, status as a protected veteran, physical or mental disability, medical condition, genetic information or characteristics (or those of a family member) or any other consideration made unlawful by applicable federal, state or local laws. We also consider qualified applicants regardless of criminal histories, consistent with legal requirements. If you require a medical or religious accommodation during the job application process, please complete this form for support. By submitting this application, I certify that all information submitted in my application and throughout the hiring process is true, accurate, and complete to the best of my knowledge. I understand that any false statement, omission, or misrepresentation may disqualify me from employment consideration or result in termination if discovered after hire.
About Zscaler Zscaler accelerates digital transformation to ensure our customers can be more agile, efficient, resilient, and secure. As an AI-forward enterprise, we are constantly pushing the envelope, leveraging the world’s largest security data lake to power our cloud-native Zero Trust Exchange platform. This innovation protects our customers from cyberattacks and data loss by securely connecting users, devices, and applications in any location. Here, impact in your role matters more than title and trust is built on results. We say, impact over activity. We seek innovators who actively use AI to amplify their impact and who thrive in an environment where we leverage intelligent systems to stay ahead of evolving threats. We believe in transparency and value constructive, honest debate—we’re focused on getting to the best ideas, faster. We build high-performing teams that can make an impact quickly and with high quality. To do this, we are building a culture of execution centered on customer obsession, collaboration, ownership, and accountability. We value high-impact, high-accountability with a sense of urgency where you’re enabled to do your best work and embrace your potential. If you’re driven by purpose, thrive on solving complex challenges, and want to be part of the team that’s helping to secure the AI age, we invite you to bring your talents to Zscaler and help shape the future of cybersecurity. Role We are looking for a Sr Staff Software Development Engineer to join our team. This is a hybrid role based in Hyderabad, reporting to the Senior Manager of Software Engineering. You will join the Engineering team that built the world’s largest cloud security platform from the ground up. As a leader in our multitenant architecture, you will bring your vision and passion to help organizations worldwide harness speed and agility with a cloud-first strategy. What you’ll do (Role Expectations) * Understand how to build and operate high scale systems * Provide service and product wide architectural guidance, and drive impactful technical decisions * Establish and enforce best practices for coding, testing, observability, and CI/CD pipelines to maintain high-quality, production-ready services * Drive cross-team collaboration and lead initiatives in performance optimization, reliability improvements, and adoption of new technologies/tools to accelerate feature velocity Who You Are (Success Profile) * You thrive in ambiguity. You're comfortable building the path as you walk it. You thrive in a dynamic environment, seeing ambiguity not as a hindrance, but as the raw material to build something meaningful. * You act like an owner. Your passion for the mission fuels your bias for action. You operate with integrity because you genuinely care about the outcome. True ownership involves leveraging dynamic range: the ability to navigate seamlessly between high-level strategy and hands-on execution. * You are a problem-solver. You love running towards the challenges because you are laser-focused on finding the solution, knowing that solving the hard problems delivers the biggest impact. * You are a high-trust collaborator. You are ambitious for the team, not just yourself. You embrace our challenge culture by giving and receiving ongoing feedback—knowing that candor delivered with clarity and respect is the truest form of teamwork and the fastest way to earn trust. * You are a learner. You have a true growth mindset and are obsessed with your own development, actively seeking feedback to become a better partner and a stronger teammate. You love what you do and you do it with purpose. What We’re Looking for (Minimum Qualifications) * Foundational understanding of AI/ML technologies and experience leveraging, securing, or positioning AI-driven solutions to optimize outcomes within your functional domain * 7+ years of experience in Java/Go coding in a highly distributed and enterprise-scale environment * Working knowledge of cloud infrastructure services on AWS or Azure * Strong experience with distributed systems and microservices architecture * Strong database knowledge of SQL constructs and data modeling * Bachelor's degree in computer science or equivalent experience What Will Make You Stand Out (Preferred Qualifications) * Experience building full CI/CD systems leveraging Kubernetes, web service frameworks, and streaming data infrastructure like Apache Spark, Kafka, Druid, or ElasticSearch * Experience building reliable data tiers using Postgres, Redis, Terraform, Ansible, graph databases like Neo4j or AWS Neptune, and developing GraphQL APIs * Deep knowledge of identity and access management systems (Okta, SAML, OAuth), test automation, Python programming, and designing fault-tolerant distributed systems for networking or cloud security products #LI-Hybrid #LI-AN4 At Zscaler, we are committed to building a team that reflects the communities we serve and the customers we work with. We foster an inclusive environment that values all backgrounds and perspectives, emphasizing collaboration and belonging. Join us in our mission to make doing business seamless and secure. Our Benefits program is one of the most important ways we support our employees. Zscaler proudly offers comprehensive and inclusive benefits to meet the diverse needs of our employees and their families throughout their life stages, including: * Various health plans * Time off plans for vacation and sick time * Parental leave options * Retirement options * Education reimbursement * In-office perks, and more! Learn more about Zscaler's hybrid working model and benefits here. By applying for this role, you adhere to applicable laws, regulations, and Zscaler policies, including those related to security and privacy standards and guidelines. Zscaler is committed to providing equal employment opportunities to all individuals. We strive to create a workplace where employees are treated with respect and have the chance to succeed. All qualified applicants will be considered for employment without regard to race, color, religion, sex (including pregnancy or related medical conditions), age, national origin, sexual orientation, gender identity or expression, genetic information, disability status, protected veteran status, or any other characteristic protected by federal, state, or local laws. See more information by clicking on the Know Your Rights: Workplace Discrimination is Illegal link. Pay Transparency Zscaler complies with all applicable federal, state, and local pay transparency rules. Zscaler is committed to providing reasonable support (called accommodations or adjustments) in our recruiting processes for candidates who are differently abled, have long term conditions, mental health conditions or sincerely held religious beliefs, or who are neurodivergent or require pregnancy-related support.