
Zscaler · Remote - California
About Zscaler Zscaler accelerates digital transformation to ensure our customers can be more agile, efficient, resilient, and secure. As an AI-forward enterpri...
About Zscaler
Zscaler accelerates digital transformation to ensure our customers can be more agile, efficient, resilient, and secure. As an
AI-forward enterprise, we are constantly pushing the envelope, leveraging the world’s largest security data lake to power our
cloud-native Zero Trust Exchange platform. This innovation protects our customers from cyberattacks and data loss by securely
connecting users, devices, and applications in any location.
Here, impact in your role matters more than title and trust is built on results. We say, impact over activity. We seek innovators
who actively use AI to amplify their impact and who thrive in an environment where we leverage intelligent systems to stay ahead
of evolving threats. We believe in transparency and value constructive, honest debate—we’re focused on getting to the best ideas,
faster. We build high-performing teams that can make an impact quickly and with high quality. To do this, we are building a
culture of execution centered on customer obsession, collaboration, ownership, and accountability.
We value high-impact, high-accountability with a sense of urgency where you’re enabled to do your best work and embrace your
potential. If you’re driven by purpose, thrive on solving complex challenges, and want to be part of the team that’s helping to
secure the AI age, we invite you to bring your talents to Zscaler and help shape the future of cybersecurity.
Role
We are looking for a Sr. Production Engineer to join our team. This role is available as a hybrid opportunity 3 days a week in San
Jose, CA or Remote reporting to Production Engineering in the Cloud Infrastructure & Operations department. Join Zscaler to be a
force multiplier for the reliability of a global platform processing 200+ billion transactions daily across tens of millions of
enterprise users.
In this role, you will provide the technical vision and hands-on execution to drive an "automation-first" culture across the
company. By maturing our observability and architectural standards, you will directly reduce our Mean Time to Mitigate (MTTM) and
shape the scalability of our globally distributed, multi-cloud infrastructure.
What you’ll do (Role Expectations)
budgets
Who You Are (Success Profile)
What We’re Looking for (Minimum Qualifications)
technical operability reviews
What Will Make You Stand Out (Preferred Qualifications)
(HAProxy), DNS at scale, and OS networking stack internals
#LI-Hybrid #LI-RT101 #LI-DY1
Zscaler’s salary ranges are benchmarked and are determined by role and level. The range displayed on each job posting reflects the
minimum and maximum target for new hire salaries for the position across all US locations and could be higher or lower based on a
multitude of factors, including job-related skills, experience, and relevant education or training.
The base salary range listed for this full-time position excludes commission/ bonus/ equity (if applicable) + benefits.
Base Pay Range
At Zscaler, we are committed to building a team that reflects the communities we serve and the customers we work with. We foster
an inclusive environment that values all backgrounds and perspectives, emphasizing collaboration and belonging. Join us in our
mission to make doing business seamless and secure.
Our Benefits program is one of the most important ways we support our employees. Zscaler proudly offers comprehensive and
inclusive benefits to meet the diverse needs of our employees and their families throughout their life stages, including:
Learn more about Zscaler's hybrid working model and benefits here.
By applying for this role, you adhere to applicable laws, regulations, and Zscaler policies, including those related to security
and privacy standards and guidelines.
Zscaler is committed to providing equal employment opportunities to all individuals. We strive to create a workplace where
employees are treated with respect and have the chance to succeed. All qualified applicants will be considered for employment
without regard to race, color, religion, sex (including pregnancy or related medical conditions), age, national origin, sexual
orientation, gender identity or expression, genetic information, disability status, protected veteran status, or any other
characteristic protected by federal, state, or local laws. See more information by clicking on the Know Your Rights: Workplace
Discrimination is Illegal link.
Pay Transparency
Zscaler complies with all applicable federal, state, and local pay transparency rules.
Zscaler is committed to providing reasonable support (called accommodations or adjustments) in our recruiting processes for
candidates who are differently abled, have long term conditions, mental health conditions or sincerely held religious beliefs, or
who are neurodivergent or require pregnancy-related support.
The impact you will have: As Staff MLOps Engineer, you will define and build Elliptic's Enterprise MLOps platform. Elliptic has growing ML capability across several teams, an established model registry, and a maturing model risk management practice. What is missing is the unified platform layer that ties training, deployment, monitoring, and governance together into a coherent, scalable discipline. You will be responsible for creating that layer. Your platform will serve four distinct internal consumers, each with different needs: * Product Engineering teams building customer-facing models and customer data analytical models, who need reproducible training pipelines, CI/CD for model deployment, and low-latency serving infrastructure * Intelligence Research building frontier intelligence collection, predictive pre-screening models, and behavioural pattern detection, who need rapid experimentation, GPU orchestration, and dataset versioning * InfoSec who own the model registry and model risk management framework today, and need the platform to close execution gaps in audit trails, drift monitoring, and compliance reporting * Operations who own BI, usage prediction, and revenue opportunity signalling, and need scheduled batch inference, BI integration, and pipeline reliability The platform you build must enforce governance with enough rigour to satisfy a regulated financial crime context, while remaining flexible enough to avoid slowing down research teams who need to iterate quickly. This is a role for someone who has built ML infrastructure from the ground up before, who understands that a platform succeeds only when it is adopted, and who is comfortable making build-vs-buy decisions that others will adopt and use for years. What you will do: * Define the target-state MLOps architecture for Elliptic, covering model training pipelines, serving infrastructure, monitoring, feature management, and governance, and produce the architecture decision records that inform investment decisions * Make and document build-vs-buy-vs-stop recommendations with clear cost modelling and trade-off analysis, evaluating vendors, open-source tools, and managed services against Elliptic's constraints (AWS-primary, Databricks ecosystem) * Work with InfoSec to improve the existing model registry and model risk management framework, closing identified gaps in metadata, lineage, approval workflows, and drift/bias detection * Build model training pipelines, CI/CD for ML, and serving infrastructure, working directly with a small group of infrastructure engineers to ship production-grade platform capabilities * Instrument observability across the ML lifecycle: training metrics, serving latency and throughput, data quality, and prediction drift, integrating with Elliptic's existing observability stack * Work directly with data scientists and ML engineers across all four consumer groups to onboard them onto the platform, writing documentation, runbooks, and reference architectures that lower the barrier to self-service You will be a great fit here if you: * Have built MLOps platforms or ML infrastructure from the ground up, and can speak to what worked, what didn't, and why * Have operated in a regulated industry (e.g. compliance, financial) and have hands on experience building ML infrastructure to meet those regulatory demands * Think about ML infrastructure the way the best platform engineers think about data infrastructure: as a set of foundations with internal customers whose needs must be understood and balanced * Are comfortable operating in ambiguity, making decisions with incomplete information, and creating structure where none exists, while remaining open to changing course when better information arrives * Influence through clarity, evidence, and the quality of your work rather than positional authority. You earn adoption by making the platform genuinely better than the alternative * Care about production engineering quality: you write production-grade code, your systems are tested, observable, documented, and designed for others to operate Our ideal candidate has: * Deep hands-on experience building MLOps platforms, including model registries, feature stores, and ML pipeline orchestration * Working knowledge of model serving patterns: real-time inference, batch prediction, A/B deployment, and deployment strategies * AWS infrastructure experience (ECS/EKS, S3, IAM, networking) and comfort operating in a Databricks ecosystem or equivalent lakehouse architecture * Experience with model monitoring: model evaluation, data drift detection, prediction drift, and performance degradation alerting * A track record of building something from zero and bringing it to a state where others could operate and extend it * Experience in a regulated industry (fintech, financial services, healthcare) where model governance is a compliance requirement * See AI as a core part of how modern engineering gets done, not a passing trend. You actively use it to think faster, prototype faster, and pressure-test your own designs, and you're excited that the bar keeps rising. * Prior experience running formal build-vs-buy evaluations with written decision records Bonus Points for: * Familiarity with model risk management frameworks and the ability to connect governance practices to regulatory expectations * Experience working simultaneously with research-oriented ML teams and production-oriented engineering teams, and understanding how their needs diverge * Infrastructure-as-code fluency (Terraform) * Experience with ClickHouse or similar OLAP engines for low-latency ML feature serving * Blockchain or crypto domain knowledge * Experience working in fraud detection and modelling * Contributions to open-source MLOps tooling JOB BENEFITS > How we work: * Hybrid working and the option to work from almost anywhere for up to 90 days per year * £500 Remote working budget to set up your home office space > Learning & Development: * $1,000 Learning & Development budget to use on anything (agreed with your manager) that contributes to your growth and development > Vacation/ Leave: * Holidays: 25 days of annual leave + bank holidays * An extra day for your birthday * Enhanced parental leave: we provide eligible employees, regardless of gender or whether they become a parent by birth or adoption, 16 weeks fully-paid leave and leave. > Benefits: * Private Health Insurance - we use Vitality! * Full access to Spill Mental Health Support * Life Assurance: we hope you will never need this - but our cover is for 4 times your salary to your beneficiaries * Cycle to Work Scheme
Multiverse is the upskilling platform for AI and Tech adoption. We have partnered with 1,500+ companies to deliver a new kind of learning that's transforming today’s workforce. Our upskilling apprenticeships are designed for people of any age and career stage to build critical AI, data, and tech skills. Our learners have driven $2bn+ ROI for their employers, using the skills they’ve learned to improve productivity and measurable performance. In April 2026, we announced $70 million in strategic funding, led by Schroders Capital, with participation from StepStone Group, Lightspeed Venture Partners and General Catalyst. At an increased valuation of $2.1bn, the round makes us Europe’s first EdTech double unicorn. But we aren’t stopping there. With a strong operational footprint and 800+ employees, we have ambitious plans to continue scaling. We’re building a world where tech skills unlock people’s potential and output. Join Multiverse and power our mission to equip the workforce to win in the AI era. THE ROLE Multiverse is the UK's largest apprenticeship provider and its first EdTech unicorn. The current state of AI presents a huge opportunity to reshape the future of education and workforce development. Multiverse is in a uniquely strong position to do that, and getting it right has implications beyond the company: for the UK tech sector and the broader economy. The AI Transformation team exists to make that real, starting with Multiverse itself. This is not a team that bolts AI onto the edges of the business or ships a handful of internal productivity tools. The mandate is bigger: to rebuild how the company actually works, function by function, and to establish the engineering practices that make Multiverse an AI-first company from the core out. That work matters twice over. Get it right inside Multiverse and we move faster, serve learners better, and operate at a level few organisations can match. But Multiverse also exists to build the workforce that every other company is reaching for. The way we transform ourselves becomes the standard we set for everyone else. You are not just changing one company, you are building the blueprint others will follow. The team is one small, focused squad, accountable for outcomes end to end. You work closely with the wider engineering org building Multiverse's customer-facing product, and alongside the teams whose work you are helping to reinvent. The structure is flat and fast. No shared queues, no bureaucratic overhead between having an idea and shipping it. Whilst we are building something entirely new, Multiverse has an established product, existing infrastructure, and engineering teams in London and Berlin. You need to be as comfortable integrating existing systems and working across team boundaries as you are building new ones from scratch. WHAT YOU WILL DO Own the architecture of our internal agentic operating system. The team's work spans the full surface of how Multiverse operates. You own the technical architecture of our agentic operating system: the agent orchestration, context strategy, tool integrations, evaluation framework, and production operation. Your design decisions shape what is possible for human and AI teams at Multiverse Ship production AI agent systems. This is a building role. You write code, review code, and own the quality of what goes to production. You will personally build and deliver significant agent systems. On a squad this size, nobody leads from a whiteboard. Design multi-agent coordination. Task decomposition across agents, handoff protocols, shared state management, orchestration logic. You know the difference between agents that genuinely coordinate and agents that run sequentially and hope for the best. You design the patterns that make multi-agent systems reliable. Build the evaluation and quality infrastructure. Automated eval pipelines, human-in-the-loop review systems, regression testing for prompt changes, domain-specific quality metrics. You treat evaluation as a first-class engineering concern and build the systems that make it possible at scale. Drive cost engineering. Token economics, caching strategies, model routing, prompt optimisation. The cost profile of production AI systems requires active engineering attention, and you build the cost awareness and tooling into the architecture rather than bolting it on later. Build the integration layer that makes existing Multiverse systems agent-accessible. APIs, MCPs, shared data contracts, and the tooling that connects agents to the platform, content systems, and the tools the company runs on. This means building real working relationships with engineering teams across London and designing interfaces that serve both sides well. Set the standard. You define patterns for prompt management, retrieval, guardrails, and testing that the wider team and eventually the whole organisation adopts — and that, in time, shape how the companies who learn from Multiverse do this too. You do this through code, documentation, and architectural decisions, not through mandates. Mentor the team. Code review, architectural guidance, pairing on the hardest problems. You are not a line manager, but your technical leadership directly shapes the growth of the engineers around you. WHAT WE ARE LOOKING FOR Production AI Agent Engineering You have shipped multi-agent systems or complex AI products to real users. You understand the engineering challenges that make agent systems a distinct discipline: * Context management. Designing what enters the context window and what stays out. Retrieval strategies, chunking, conversation memory, summarisation, and the cost/quality trade-offs of each. You have made these decisions in production and seen the consequences. * Model selection and routing. Choosing the right model for each task based on capability, latency, cost, and reliability. Building routing logic that matches work to the appropriate model rather than defaulting to one. * Cost engineering. Token economics, caching, prompt optimisation, batching. You know the difference between a prototype that works and a production system that works at sustainable cost. You have built systems where cost was an engineering constraint, not someone else's problem. * Tool use and agent augmentation. Designing what capabilities agents can reach: tool descriptions that models use reliably, failure handling, MCPs or equivalent interfaces. You understand that the quality of the tool layer determines whether agents are useful or fragile. * Multi-agent coordination. Task decomposition across agents, handoff protocols, shared state, orchestration logic. You have built systems where multiple agents work together within a product domain and understand the architectural patterns that make coordination reliable. * Evaluation and quality. Building eval frameworks for AI output: accuracy, helpfulness, safety, domain-specific criteria. Automated pipelines and human-in-the-loop review. You would not ship an agent system without a quality baseline. Product Thinking and Entrepreneurial Instinct On a small squad there is no gap between product thinking and engineering. You own the problem from user need to production system. You can sit with the people whose work you are transforming, understand their workflow, identify the highest-value intervention, and build it without waiting for a product manager to write a spec. You have either built something yourself (a product, a startup, a project with real users) or operated with that founder mindset inside a larger organisation. You understand that speed matters and that shipping something useful beats polishing something theoretical. AI-Native Engineering You build with Claude Code daily. You set context and constraints before generating code. You review AI output critically. You augment the tool with skills, system prompts, and domain context to make it effective. This is how the team works, and you help define what good looks like. Full-Stack Delivery You work across the stack: LLM integration, backend services, data pipelines, and enough frontend to ship end to end. The boundaries between these layers dissolve in agent systems, and so should your willingness to work across them. Communication You can explain technical strategy to a CPO, walk a product manager through a cost trade-off, and give direct feedback in code review. You represent the team's technical approach in cross-functional forums with product, design, learning design, compliance, and other engineering teams. You document decisions, not just code. WHAT WOULD SET YOU APART * Experience in EdTech, regulated content, or domains where AI output quality has compliance or accreditation implications * Background as a founding engineer or technical co-founder * Published thinking or external contributions in AI engineering (talks, writing, open source) * Experience designing platform layers that other teams build on * Practical experience with MCP (Model Context Protocol) or equivalent agent integration standards WHAT WE ARE NOT LOOKING FOR * Pure ML research without production engineering experience. We need builders * Narrow specialism. This team works across the full stack of an AI product. If you only do infrastructure, or only do model training, or only do frontend, this is the wrong fit * People who need a detailed spec, a sprint plan, and a standup before they can write a line of code. We ship fast and iterate * Candidates whose experience is limited to wrapping LLM APIs in thin application layers. We need depth in agent architecture, context strategy, tool design, and multi-agent coordination * Engineers who optimise for technical elegance over user outcomes. The architecture serves the product Benefits * Time off - 27 days holiday, plus 5 additional days off: 1 life event day, 2 volunteer days, 2 company-wide wellbeing days (M-Powered Weekend) and 8 bank holidays per year * Health & Wellness- private medical Insurance with Bupa, a medical cashback scheme, life insurance, gym membership & wellness resources through Wellhub and access to Spill - all in one mental health support * Hybrid work offering - for most roles we collaborate in the office three days per week with the exception of Coaches and Instructors who collaborate in the office once a month * Work-from-anywhere scheme - you'll have the opportunity to work from anywhere, up to 10 days per year * Space to connect: Beyond the desk, we make time for weekly catch-ups, seasonal celebrations, and have a kitchen that’s always stocked! Our Commitment to Diversity, Equity and Inclusion We’re an equal opportunities employer. And proud of it. Every applicant and employee is afforded the same opportunities regardless of race, colour, ancestry, religion, sex, national origin, sexual orientation, age, citizenship, marital status, disability, gender, gender identity or expression, or veteran status. This will never change. Read our Equality, Diversity & Inclusion policy here. Our Commitment to Safeguarding Multiverse is committed to safeguarding and promoting the welfare of our learners. We expect all employees to share this commitment and adhere to our Safeguarding Policy, our Prevent Policy and all other Multiverse company policies. Successful applicants will be required to undertake at least a Basic check via the Disclosure Barring Service (DBS). For roles that will involve a Regulated Activity, successful applicants must also undergo an Enhanced DBS check, including a Children’s Barred List check and a Prohibition Order check. Roles involving Regulated Activity may interact with vulnerable groups, therefore are exempt from the Rehabilitation of Offenders Act 1974 meaning applicants are required to declare any convictions, cautions, reprimands, and final warnings. Providing false information is an offence and could result in the application being rejected or summary dismissal if the applicant has been selected, and possible referral to the police and the DBS.
Company Description At Amwell, we’re transforming healthcare for all—powered by technology and inspired by people. Here, your ideas don’t just matter—they drive real change, improving lives on a global scale. We marry technology and innovation with clinical excellence to provide trusted solutions that solve the healthcare industry’s biggest pain points and are on a mission to enable greater access to more convenient, affordable, and effective care. We do this through our technology-enabled care platform that is designed to help our clients achieve their digital care ambitions – today and in the future. We offer programs spanning the full care continuum, including urgent, acute and specialty care, behavioral health, and services for the treatment of chronic conditions such as heart and cardiometabolic diseases. Programs are powered by Amwell as well as our growing partner network. For almost two decades, Amwell has proudly served some of the largest and most sophisticated healthcare organizations in the U.S. and worldwide. Our team is passionate about technology’s role in transforming care delivery and making it more equitable, accessible, efficient, cost-effective and navigable for all. Brief Overview As a Staff Site Reliability Engineer (P4), you will define and elevate the reliability standards across the platform. This role goes beyond owning individual services — you will establish the patterns, practices, and tooling that enable all teams to build and operate reliable systems at scale. You will operate across team boundaries, identifying systemic reliability risks and designing cross-cutting solutions that improve the overall health of the platform. Acting as a bridge between service-level reliability and organizational maturity, you will help ensure reliability becomes a built-in property of the system rather than a reactive effort. This role combines deep technical expertise with strong leadership and influence. You will mentor senior engineers, guide architectural decisions, and promote a culture of proactive reliability, observability, and operational excellence across the organization. Core Responsibilities * Define and evolve reliability standards, patterns, and tooling adopted across the platform. * Own the reliability posture for critical service domains and drive architectural reviews to ensure reliability, operability, and recovery are first-class concerns. * Design and implement cross-cutting reliability mechanisms such as circuit breakers, retry policies, graceful degradation, and load shedding. * Establish and maintain scalable SLO frameworks that teams can adopt with minimal friction. * Lead complex, multi-service incident response as an incident commander and drive high-quality postmortems focused on systemic improvements. * Identify recurring incident patterns and implement structural solutions to prevent future failures. * Improve incident response processes, tooling, escalation paths, and communication practices. * Design and drive observability strategies across services, including metrics, logs, traces, and alerting systems. * Ensure alerting is actionable and aligned with SLOs, and build shared dashboards and runbooks to reduce time to resolution. * Collaborate with Platform Engineering to strengthen infrastructure reliability across Kubernetes (EKS), networking, and data systems. * Contribute to infrastructure as code for reliability-critical components and validate disaster recovery, backup, and restore strategies. * Drive chaos engineering practices and ensure deployment pipelines include reliability safeguards such as health checks, canary releases, and rollback automation. * Lead capacity planning and performance optimization efforts across services and shared infrastructure. * Identify bottlenecks and failure risks across distributed systems and design solutions that improve resilience and recovery. * Mentor engineers through design reviews, incident leadership, and knowledge sharing. * Promote best practices, improve operational maturity across teams, and influence engineering culture toward proactive reliability investments. * Represent reliability concerns in cross-functional planning and contribute to long-term platform strategy. Qualifications * 8+ years of experience in Site Reliability Engineering, infrastructure, or production engineering roles. * Strong experience operating and improving large-scale production systems in AWS environments. * Deep expertise in Kubernetes (preferably EKS), including networking, scheduling, and observability. * Hands-on experience with Infrastructure as Code tools such as Terraform or CDKTF. * Advanced understanding of distributed systems, networking, and failure modes. * Experience designing and managing observability stacks (e.g., Prometheus, Grafana, OpenSearch, OpenTelemetry). * Proven experience leading incident response for complex, multi-service production environments. * Demonstrated ability to drive systemic reliability improvements across teams and platforms. * Strong written communication skills, including experience creating postmortems, design documents, runbooks, and architectural proposals. * Experience with service mesh technologies (e.g., Istio) and mTLS is a plus. * Familiarity with GitOps workflows (e.g., ArgoCD, Flux) is a plus. * Experience working in compliance-driven environments (e.g., HIPAA, SOC2, FedRAMP) is preferred. * Exposure to chaos engineering practices and cost-aware infrastructure design (FinOps) is a plus. Do Well. Live Well. At Amwell. Driven by our mission and values, we foster a workplace where Delivering Awesome, being Customer First and operating as One Team aren’t just aspirations – they are how we work, every day. Our people are our greatest asset. We strive to empower their growth and development not only as Amwellians but as individuals, through generous total rewards packages, a virtual-first work environment, work-life flexibility, including Summer Fridays and designated Mental Health Days, as well as opportunities to stretch and learn – to name a few. It’s our people who truly differentiate us. Ask anyone and they’ll tell you – you’ll never work with more passionate, more driven and more caring team members. We champion a culture of respect and inclusion, accountability and integrity, innovation and collaboration. At Amwell, you’ll do the most meaningful work of your career—improving healthcare for millions, growing alongside incredible teammates, and being valued for who you are. Working at Amwell: Amwell is changing how care is delivered through online and mobile technology. We strive to make the hard work of healthcare look easy. In order to make this a reality, we look for people with a fast-paced, mission-driven mentality. We’re a culture that prides itself on quality, efficiency, smarts, initiative, creative thinking, and a strong work ethic. Our Core Values include One Team, Customer First, and Deliver Awesome. Customer First and Deliver Awesome are all about our product and services and how we strive to serve. As part of One Team, we operate the Amwell Cares program, which brings needed assistance to our communities, whether that be free healthcare for the underserved or for people affected by natural disasters, support for equality, honoring doctors and nurses, or annual Amwell-matched donations to food banks. Amwell aims to be a force for good for our employees, our clients, and our communities. Amwell cares deeply about and supports Diversity, Equity and Inclusion. These initiatives are highlighted and reflected within our Three DE&I Pillars - our Workplace, our Workforce and our Community! Benefits Additional Benefits * Medical Plan Coverage provided by Colmédica * Plan Coverage provided by Pan American * Hybrid Allowance * Additional Paid Time Off * Maternity Leave 18 weeks * Parental/Paternity Leave 2 mandatory weeks + 4 weeks * Mental Health and Resiliency * Virtual Second Opinion with the Cleveland Clinic Coverage * LinkedIn Learning * Rewards and Recognition * Service Anniversaries * Annual Bonus * Referral Program * Amwell tuition reimbursement benefit https://business.amwell.com/privacy-policy/ Privacy Notice