
Okta · Washington
Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure t...
Secure Every Identity, from AI to Human
Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables
organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world
stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence.
This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk.
Okta is the leading independent identity partner, securely connecting the right people to the right technologies at the right
time. With over 20,000 customers and a billion registered users, we provide the foundation for secure enterprise management and
seamless user experiences. As a trusted U.S. Federal vendor, we maintain FedRAMP High and IL4 compliance to support the
government’s most critical missions.
Our team spans the globe and Okta’s platform to build, deliver, and maintain Okta’s legendary resiliency and reliability. We’re
the SME’s for a vast number of synchronous and asynchronous workloads, data stores, and deployments, running the majority of our
customer-facing Identity-as-a-Service platform. You’ll act as an advocate for reliability and a guide for global scale, ensuring
our product is always available, always secure, and Always On.
We are seeking a Staff Site Reliability Engineer (TS/SCI) to join our high-stakes National Security team. Based in the Washington,
D.C. area, with on-site customer travel, you will serve as a technical leader, delivering secure, air-gapped solutions for our
federal partners.
Security Requirement: Must be able to obtain and maintain a U.S. security clearance (Secret or Top Secret) to the extent required
by U.S. Government contracts.
The selected candidate may be subject to drug testing to the extent required by U.S. Government contracts.
adapting our existing deployments for secure federal air-gapped environments.
and implementing permanent preventive solutions.
and deep technical expertise.
secure, enterprise-grade solutions.
experience developing and troubleshooting web services on Kubernetes or similar orchestration layers.
large-scale batch processing – ideally using data warehouse products such as Snowflake, Redshift, or Databricks.
rigorous software engineering best practices.
Industry Experience: Prior experience supporting or building mission-critical Enterprise SaaS platforms.
#LI-Hybrid
Below is the annual base salary range for candidates located in California (excluding San Francisco Bay Area), Colorado, Illinois,
New York and Washington. Your actual base salary will depend on factors such as your skills, qualifications, experience, and work
location. In addition, Okta offers equity (where applicable), bonus, and benefits, including health, dental and vision insurance,
401(k), flexible spending account, and paid leave (including PTO and parental leave) in accordance with our applicable plans and
policies. To learn more about our Total Rewards program please visit: https://rewards.okta.com/us.
The annual base salary range for this position for candidates located in California (excluding San Francisco Bay Area), Colorado,
The Okta Experience
We are intentional about connection. Our global community, spanning over 20 offices worldwide, is united by a drive to innovate.
Your journey begins with an immersive, in-person onboarding experience designed to accelerate your impact and connect you to our
mission and team from day one.
Okta is an Equal Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race,
color, religion, sex, sexual orientation, gender identity, national origin, ancestry, marital status, age, physical or mental
disability, or status as a protected veteran. We also consider for employment qualified applicants with arrest and convictions
records, consistent with applicable laws.
If reasonable accommodation is needed to complete any part of the job application, interview process, or onboarding please use
this Form to request an accommodation.
Notice for New York City Applicants & Employees: Okta may use Automated Employment Decision Tools (AEDT), as defined by New York
City Local Law 144, that use artificial intelligence, machine learning, or other automated processes to assist in our recruitment
and hiring process. In accordance with NYC Local Law 144, if you are an applicant or employee residing in New York City,
please click here to view our full NYC AEDT Notice.
Reddit is a community of communities. It’s built on shared interests, passion, and trust, and is home to the most open and authentic conversations on the internet. Every day, Reddit users submit, vote, and comment on the topics they care most about. With 100,000+ active communities and approximately 126 million daily active unique visitors, Reddit is one of the internet’s largest sources of information. For more information, visit www.redditinc.com. As Reddit continues to scale globally, reliability and performance are more critical than ever. The Site Experience SRE team sits at the intersection of infrastructure, product engineering, and user experience - ensuring that every interaction across web, mobile, APIs, feeds, media delivery, and real time systems is fast, reliable, and resilient. We are looking for a Staff Site Reliability Engineer to lead reliability engineering initiatives for critical user facing systems at internet scale. In this role, you will partner closely with product and infrastructure teams to improve availability, latency, scalability, and operational excellence across Reddit’s most business critical experiences. This is a highly technical leadership role for someone who thrives in large-scale distributed systems, enjoys solving complex reliability challenges, and can influence engineering culture across the organization. WHAT YOU’LL DO: * Lead Reliability Engineering for User Experience * Drive reliability, scalability, and operational excellence for critical user facing systems and services. Improve performance and resiliency across APIs, content delivery, feed generation, search, messaging, and real-time experiences. * Architect for Scale * Partner with product and infrastructure engineering teams to design systems that remain highly available and performant under massive global load. Guide architectural decisions around failover, redundancy, graceful degradation, traffic management, and capacity planning. * Reduce Operational Risk * Identify systemic risks and reliability bottlenecks across services, dependencies, deployments, and infrastructure. Build proactive mitigation strategies and drive engineering improvements that reduce incidents and improve service health. * Drive Automation * Eliminate repetitive operational work through automation and tooling. Build systems that improve deployment safety, incident response, remediation workflows, and reliability guardrails * Incident Management * Lead complex incident response efforts across engineering teams. Drive blameless postmortems, identify root causes, and ensure sustainable long-term fixes are implemented. Influence Engineering Standards * Define and champion best practices around reliability engineering, SLIs/SLOs, capacity management, release engineering, and operational maturity across the company. Mentor and Multiply Impact * Provide technical leadership and mentorship to engineers across SRE and software engineering teams. Help shape reliability culture and raise the operational excellence bar across the organization. WHAT WE’RE LOOKING FOR * 8+ years of experience in Site Reliability Engineering, Infrastructure Engineering, or related roles operating large scale distributed systems. * Strong collaboration and communication skills with the ability to influence technical direction across teams. * Strong experience supporting high traffic, user facing production environments. * Deep understanding of one or more: distributed systems, networking, Linux systems, cloud native architectures. * Experience designing highly available systems with strong operational and reliability practices. * Strong programming skills in languages such as Go, Python, or similar. * Strong understanding of observability systems including metrics, logging, tracing, and alerting. * Experience improving reliability through SLOs, automation, incident management, and performance optimization. * Demonstrated ability to troubleshoot complex issues across applications, infrastructure, networking, and services. Nice to Have * Experience operating systems at internet scale traffic volumes. * Experience with Kubernetes, containers, cloud infrastructure, and modern deployment platforms. * Familiarity with technologies such as Prometheus, Grafana, OpenTelemetry, Envoy, Kafka, ClickHouse, Cassandra, Redis, or similar distributed infrastructure technologies. * Experience with CDN optimization, edge reliability, traffic engineering, or global infrastructure. * Contributions to open source software or participation in technical communities. * Experience leading large scale incident response and operational transformation initiatives. Why Join Reddit? You’ll help shape the reliability and performance of one of the internet’s largest platforms, influencing experiences used by millions of people every day. This is an opportunity to solve deeply complex engineering problems at massive scale while helping define the future of reliability engineering for a modern consumer platform. Benefits * Global Benefit programs that fit your lifestyle, from workspace to professional development to caregiving support * Family Planning Support * Gender-Affirming Care * Mental Health & Coaching Benefits * Group Personal Pension Scheme with Employer match * Private Medical and Dental Scheme * Income Replacement Programs * Bike to Work scheme * Flexible Vacation & Paid Volunteer Time Off * Generous Paid Parental Leave * * Generous Paid Parental Leave * * Generous Paid Parental Leave * In select roles and locations, the interviews will be recorded, transcribed and summarized by artificial intelligence (AI). You will have the opportunity to opt out of recording, transcription and summarization prior to any scheduled interviews. During the interview, we will collect the following categories of personal information: Identifiers, Professional and Employment-Related Information, Sensory Information (audio/video recording), and any other categories of personal information you choose to share with us. We will use this information to evaluate your application for employment or an independent contractor role, as applicable. We will not sell your personal information or disclose it to any third party for their marketing purposes. We will delete any recording of your interview promptly after making a hiring decision. For more information about how we will handle your personal information, including our retention of it, please refer to our Candidate Privacy Policy for Potential Employees and Contractors. Reddit is proud to be an equal opportunity employer, and is committed to building a workforce representative of the diverse communities we serve. Reddit is committed to providing reasonable accommodations for qualified individuals with disabilities and disabled veterans in our job application procedures. If, due to a disability, you need an accommodation during the interview process, please let your recruiter know.
Reddit is a community of communities. It’s built on shared interests, passion, and trust, and is home to the most open and authentic conversations on the internet. Every day, Reddit users submit, vote, and comment on the topics they care most about. With 100,000+ active communities and approximately 126 million daily active unique visitors, Reddit is one of the internet’s largest sources of information. For more information, visit www.redditinc.com. As Reddit continues to scale globally, reliability and performance are more critical than ever. The Site Experience SRE team sits at the intersection of infrastructure, product engineering, and user experience - ensuring that every interaction across web, mobile, APIs, feeds, media delivery, and real time systems is fast, reliable, and resilient. We are looking for a Staff Site Reliability Engineer to lead reliability engineering initiatives for critical user facing systems at internet scale. In this role, you will partner closely with product and infrastructure teams to improve availability, latency, scalability, and operational excellence across Reddit’s most business critical experiences. This is a highly technical leadership role for someone who thrives in large-scale distributed systems, enjoys solving complex reliability challenges, and can influence engineering culture across the organization. WHAT YOU’LL DO: * Lead Reliability Engineering for User Experience * Drive reliability, scalability, and operational excellence for critical user facing systems and services. Improve performance and resiliency across APIs, content delivery, feed generation, search, messaging, and real-time experiences. * Architect for Scale * Partner with product and infrastructure engineering teams to design systems that remain highly available and performant under massive global load. Guide architectural decisions around failover, redundancy, graceful degradation, traffic management, and capacity planning. * Reduce Operational Risk * Identify systemic risks and reliability bottlenecks across services, dependencies, deployments, and infrastructure. Build proactive mitigation strategies and drive engineering improvements that reduce incidents and improve service health. * Drive Automation * Eliminate repetitive operational work through automation and tooling. Build systems that improve deployment safety, incident response, remediation workflows, and reliability guardrails * Incident Management * Lead complex incident response efforts across engineering teams. Drive blameless postmortems, identify root causes, and ensure sustainable long-term fixes are implemented. Influence Engineering Standards * Define and champion best practices around reliability engineering, SLIs/SLOs, capacity management, release engineering, and operational maturity across the company. Mentor and Multiply Impact * Provide technical leadership and mentorship to engineers across SRE and software engineering teams. Help shape reliability culture and raise the operational excellence bar across the organization. WHAT WE’RE LOOKING FOR * 8+ years of experience in Site Reliability Engineering, Infrastructure Engineering, or related roles operating large scale distributed systems. * Strong collaboration and communication skills with the ability to influence technical direction across teams. * Strong experience supporting high traffic, user facing production environments. * Deep understanding of one or more: distributed systems, networking, Linux systems, cloud native architectures. * Experience designing highly available systems with strong operational and reliability practices. * Strong programming skills in languages such as Go, Python, or similar. * Strong understanding of observability systems including metrics, logging, tracing, and alerting. * Experience improving reliability through SLOs, automation, incident management, and performance optimization. * Demonstrated ability to troubleshoot complex issues across applications, infrastructure, networking, and services. Nice to Have * Experience operating systems at internet scale traffic volumes. * Experience with Kubernetes, containers, cloud infrastructure, and modern deployment platforms. * Familiarity with technologies such as Prometheus, Grafana, OpenTelemetry, Envoy, Kafka, ClickHouse, Cassandra, Redis, or similar distributed infrastructure technologies. * Experience with CDN optimization, edge reliability, traffic engineering, or global infrastructure. * Contributions to open source software or participation in technical communities. * Experience leading large scale incident response and operational transformation initiatives. Why Join Reddit? You’ll help shape the reliability and performance of one of the internet’s largest platforms, influencing experiences used by millions of people every day. This is an opportunity to solve deeply complex engineering problems at massive scale while helping define the future of reliability engineering for a modern consumer platform. Benefits * Global Benefit programs that fit your lifestyle, from workspace to professional development to caregiving support * Family Planning Support * Gender-Affirming Care * Mental Health & Coaching Benefits * Private Medical, Dental, and Vision Benefits * Personal Retirement Savings Account with matching contribution * Cycle to Work and Tax Saver schemes * Flexible Vacation & Paid Volunteer Time Off * Generous Paid Parental Leave In select roles and locations, the interviews will be recorded, transcribed and summarized by artificial intelligence (AI). You will have the opportunity to opt out of recording, transcription and summarization prior to any scheduled interviews. During the interview, we will collect the following categories of personal information: Identifiers, Professional and Employment-Related Information, Sensory Information (audio/video recording), and any other categories of personal information you choose to share with us. We will use this information to evaluate your application for employment or an independent contractor role, as applicable. We will not sell your personal information or disclose it to any third party for their marketing purposes. We will delete any recording of your interview promptly after making a hiring decision. For more information about how we will handle your personal information, including our retention of it, please refer to our Candidate Privacy Policy for Potential Employees and Contractors. Reddit is proud to be an equal opportunity employer, and is committed to building a workforce representative of the diverse communities we serve. Reddit is committed to providing reasonable accommodations for qualified individuals with disabilities and disabled veterans in our job application procedures. If, due to a disability, you need an accommodation during the interview process, please let your recruiter know.
Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. Okta’s Workforce Identity Cloud Security Engineering group is looking for an experienced and passionate Staff Site Reliability Engineer to join a team focused on designing and developing Security solutions to harden our cloud infrastructure. We embrace innovation and pave the way to transform bright ideas into excellent security solutions that help run large-scale, critical infrastructure. We encourage you to prescribe defense-in-depth measures, industry security standards and enforce the principle of least privilege to help take our Security posture to the next level. Our Infrastructure Security team has a niche skill-set that balances Security domain expertise with the ability to design, implement, rollout infrastructure across multiple cloud environments without adding friction to product functionality or performance. We are responsible for the ever-growing need to improve our customer safety and privacy by providing security services that are coupled with the core Okta product. This is a high-impact role in a security-centric, fast-paced organization that is poised for massive growth and success. You will act as a liaison between the Security org and the Engineering org to build technical leverage and influence the security roadmap. You will focus on engineering security aspects of the systems used across our services. Join us and be part of a company that is about to change the cloud computing landscape forever. As a Staff Engineer, you should be able to identify gaps, propose innovative solutions, and contribute to roadmaps while driving alignment across multiple teams within the organization. Additionally, you should serve as a role model, providing technical mentorship to junior team members and fostering a culture of learning and growth What are we looking for? We are looking for a security-first SRE engineer who doesn't just "flag" issues but builds the automation to solve them. You should have a deep-seated intuition for cloud-native security and a proven track record of hardening large-scale GCP and AWS environments. As a Technical SME, you will design and build production infrastructure with a "security-at-scale" mindset. What You Will Work On? Security Evangelism: Lead initiatives to strengthen our security posture for critical infrastructure and promote best practices across the engineering organization. Incident Response & Reliability: Respond to production security incidents, perform root cause analysis, and build automated preventions to ensure high performance and reliability. Automated Hardening: Identify manual security processes and automate them using custom tooling and CI/CD integrations. Architecture & Documentation: Develop technical documentation, runbooks, and procedures for a 24x7 online environment. Platform Evolution: Continuously evolve our monitoring platforms, moving from simple auditing to active, automated prevention. Minimum Required Knowledge, Skills, & Abilities: Experience: 8+ years of experience architecting and running complex cloud networking and infrastructure, with at least 7+ years specialized in DevSecOps or Cloud Security. GCP Expertise: Minimum 3+ years of deep, hands-on experience securing GCP (GKE, GCE, Shared VPC etc). Infrastructure as Code (IaC): 10+ years of experience using Terraform and Chef to manage complex cloud resources and OS hardening. Automation Mastery: Expert-level proficiency in Go, Python, or Ruby for building custom security tooling and automated remediation. Hardened Containers: Proven track record of securing containerized workloads, including image scanning, K8s RBAC, and runtime security tools (e.g., CrowdStrike Falcon, Falco, or gVisor). Unflappable Troubleshooting: A "see a problem, fix the problem" mindset with the ability to debug complex networking, IAM, or performance issues under pressure. Security Foundations: Strong grasp of Linux internals, OS hardening (CIS benchmarks), and IP protocols (TLS/SSL, DNSSEC, BGP). Education: BS in Computer Science or equivalent professional experience. Key Responsibilities: IAM & Secrets Management: Design and maintain large-scale production IAM policies and secrets management workflows . Infrastructure Hardening: Implement and maintain Public Key Infrastructure (PKI) and ensure all GCE/GKE environments meet strict compliance standards. Operational Excellence: Utilize industry-standard tools like OSQuery, Splunk, Chronicle, Nessus, or Qualys/ Crowdstrike to monitor system health and security telemetry. Strategic Rollouts: Lead the phased transition of security policies from Audit/Detection mode to Blocking/Prevention mode, ensuring zero impact on production uptime. Bonus Points For: Multi-Cloud IAM Governance: Experience designing a unified IAM framework across AWS and GCP, utilizing federated Identities such as Workload, Workforce Identity Federation with understanding of SAML & OIDC auth mechanism and automated "Least Privilege" enforcement. Cloud-Native Reliability Engineering: Deep understanding of multi-cloud reliability patterns, maintaining high availability (HA) during security patching or infrastructure-wide hardening. Hardened Kubernetes Orchestration: Advanced experience securing GKE, EKS, and kOps, specifically implementing Pod Security Standards, Network Policies, and Admission Controllers for a "Zero-Trust" posture. Threat Modeling:Security Reviews & Threat Modeling at both Design & Implementation scope. The Okta Experience * Supporting Your Well-Being * Driving Social Impact * Developing Talent and Fostering Connection + Community We are intentional about connection. Our global community, spanning over 20 offices worldwide, is united by a drive to innovate. Your journey begins with an immersive, in-person onboarding experience designed to accelerate your impact and connect you to our mission and team from day one. Okta is an Equal Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, ancestry, marital status, age, physical or mental disability, or status as a protected veteran. We also consider for employment qualified applicants with arrest and convictions records, consistent with applicable laws. If reasonable accommodation is needed to complete any part of the job application, interview process, or onboarding please use this Form to request an accommodation. Notice for New York City Applicants & Employees: Okta may use Automated Employment Decision Tools (AEDT), as defined by New York City Local Law 144, that use artificial intelligence, machine learning, or other automated processes to assist in our recruitment and hiring process. In accordance with NYC Local Law 144, if you are an applicant or employee residing in New York City, please click here to view our full NYC AEDT Notice.