
Northflank · Remote
Northflank is a cutting-edge cloud platform enabling developers to build and ship highly scalable, full-stack applications faster than ever before. We are a ven...
Northflank is a cutting-edge cloud platform enabling developers to build and ship highly scalable, full-stack applications faster
than ever before. We are a venture-backed company, and our platform is used by tens of thousands of developers worldwide in
production.
We're seeking a talented Cloud Infrastructure Engineer to join our team of passionate professionals. In this role, you'll be
instrumental in architecting and maintaining the robust cloud infrastructure that powers our platform.
This position is remote.
We offer a competitive salary, equity, and benefits package, and the opportunity to be part of a fast-growing startup at the
forefront of the cloud-native revolution.
If you are passionate about infrastructure software and want to join a dynamic team, we’d love to hear from you!
Compensation
The compensation for this role will depend on a number of factors including but not limited to: skill set, experience, training,
and location.
Diversity Statement
At Northflank, we are committed to fostering a culture where every individual feels valued, respected, and supported, regardless
of their background or identity. We are dedicated to providing equal treatment and opportunity throughout hiring, selection, and
employment regardless of gender, race, religion, national origin, ethnicity, disability, gender identity/expression, sexual
orientation, veteran or military status, or any other legally protected category. Northflank is an equal opportunity employer.
Senior Infrastructure Engineer (OpenStack) Location: UK (Remote) Department: Infrastructure Reporting to: Head of Infrastructure ABOUT NEXGEN CLOUD: NexGen Cloud is the company behind Hyperstack, a full-stack AI cloud serving tens of thousands of customers from AI researchers to enterprises running the world's most compute-intensive workloads. We deliver on-demand and private GPU infrastructure to teams who treat performance as a requirement, not a feature. We're a tight-knit, fast-moving team working at the cutting edge of AI cloud infrastructure. We practice what we preach, equipping our people with AI at every level so we can solve harder problems, ship faster, and keep raising the bar for what enterprise GPU infrastructure looks like. THE ROLE: SENIOR INFRASTRUCTURE ENGINEER (OPENSTACK) This role exists because our platform is scaling quickly — and complexity comes with it. As we expand our OpenStack and Kubernetes environments globally, we need engineers who can take real ownership of how the platform is designed, operated, and improved. You'll have direct ownership over business-critical infrastructure that impacts performance, reliability, and customer experience. This is not a maintenance role. If you like solving hard problems, owning systems end-to-end, and seeing the impact of your work immediately — you'll enjoy this. WHAT YOU'LL BE DOING: Rather than a long checklist, here's what success in this role looks like: * Own the design, deployment, and operation of OpenStack and Kubernetes environments — ensuring platform performance, scalability, and resilience for GPU workloads * Build and improve infrastructure using infrastructure-as-code and GitOps practices, driving automation across provisioning, deployment, and operational workflows * Optimise GPU workload scheduling using Kubernetes and NVIDIA tooling, and implement monitoring, logging, and alerting to ensure platform stability * Lead incident response and drive continuous improvement of reliability across the platform * Maintain strong security controls across infrastructure and container layers — RBAC, network policies, and tenant isolation * Work closely with Platform, DevOps, AI, Product, and Support teams to align infrastructure capabilities with customer and platform requirements ABOUT YOU: We're more interested in how you think and work than in a perfect CV. You'll likely bring a combination of the following: ESSENTIAL * Exposure to GPU infrastructure, HPC, or large-scale compute environments * Proven experience operating Kubernetes at scale — ideally bare-metal or private cloud * Solid understanding of Linux, networking, and storage systems * Experience with infrastructure automation, CI/CD, and Git-based workflows * Strong ownership mindset — comfortable operating without heavy oversight and able to simplify and scale systems in a fast-moving environment NICE TO HAVE * Experience integrating Kubernetes with OpenStack * Ideally hands-on experience running OpenStack in production environments * Familiarity with advanced networking or cloud-native ecosystems * Contributions to open-source projects WHAT WE OFFER: * Competitive salary and annual discretionary bonus scheme * Employee wellbeing benefits * 25 days of holiday, plus public holidays * Flexible working arrangements (remote or hybrid, depending on role and location) * Real ownership and autonomy, with the trust to take initiative and experiment * The opportunity to make a visible, meaningful impact as we scale * Clear career progression and growth opportunities in a fast-growing company * A collaborative, international culture built on trust, transparency, and ownership * The chance to help shape NexGen Cloud's team, culture, and future alongside ambitious, mission-driven colleagues MORE INFORMATION Head over to our NexGen Cloud careers page to view current openings and follow us on LinkedIn and X to learn more about our journey, newest releases and hear exciting news in the neocloud space.
Who we are Moniepoint Inc. is Africa’s all-in-one financial platform, helping 20 million businesses and individuals access seamless payments, banking, credit, cross-border, and business management tools each month. As Nigeria’s largest merchant acquirer, we power most of the country’s point-of-sale (POS) transactions. Through our subsidiaries, Moniepoint Inc. processes over $250 billion in digital payment transaction value annually. About the role Engineering at Moniepoint is an inspired, customer-focused community dedicated to crafting solutions that redefine our industry. Our infrastructure runs on some of the cool tools that excite infrastructure engineers - kubernetes, docker etc. We also make business decisions based on the large stream of data we receive daily, so we work daily with big data, perform data analytics and build models to make sense of the noise and give our customers the best experience. Curious about what makes Moniepoint an incredible place to work? Check out posts on how we cultivate a culture of innovation, teamwork, and growth. Position Overview We are seeking an experienced Cloud Engineer to design, implement, and manage our multi-cloud infrastructure. The ideal candidate will have deep expertise in cloud platforms, container orchestration, infrastructure automation, CI/CD pipelines, and observability solutions, ensuring scalable, reliable, and cost-effective cloud operations across multiple cloud providers. Principal Duties and Responsibilities Cloud Infrastructure Management * Design, deploy, and manage multi-cloud infrastructure across Google Cloud Platform (GCP), Amazon Web Services (AWS), Azure, and Oracle Cloud Infrastructure (OCI) * Architect and implement highly available, fault-tolerant, and scalable cloud solutions * Manage cloud resources including compute instances and networking components * Design and implement disaster recovery and business continuity plans for cloud workloads * Migrate on-premises applications and services to cloud environments with minimal disruption * Optimize cloud resource utilization and implement auto-scaling policies * Maintain comprehensive documentation of cloud architectures, configurations, and runbooks Kubernetes & Container Orchestration * Design, deploy, and manage production-grade Kubernetes clusters across multiple cloud providers * Implement and maintain container orchestration strategies for microservices architectures * Configure and manage Kubernetes resource objects * Manage Kubernetes cluster upgrades, scaling, and performance optimization * Troubleshoot complex container and orchestration issues in production environments * Implement multi-cluster and multi-region Kubernetes deployments for high availability Service Mesh & Advanced Networking * Design, deploy, and manage Istio service mesh for microservices communication and observability * Configure Istio traffic management, including virtual services, destination rules, and gateways * Implement advanced traffic routing (canary deployments, A/B testing, traffic splitting) using Istio * Deploy and manage Istio observability components (telemetry, distributed tracing, service graphs) * Implement circuit breaking, retries, timeouts, and fault injection for resilience testing * Configure Istio ingress and egress gateways for external traffic management * Monitor and optimize service mesh performance and resource utilization * Implement multi-cluster service mesh architectures across different cloud providers Reverse Proxy & Load Balancing * Deploy, configure, and manage HAProxy for high-performance load balancing and reverse proxy * Implement HAProxy ACLs, backend routing, health checks, and session persistence * Design and implement Nginx as reverse proxy for web applications and API gateways * Configure Nginx for rate limiting and request filtering * Implement Nginx load balancing algorithms and upstream health monitoring * Manage Nginx Plus features for advanced traffic management and monitoring * Optimize HAProxy and Nginx performance for high-throughput environments Infrastructure as Code & Configuration Management * Develop and maintain infrastructure as code using Terraform * Create reusable, modular Terraform configurations for various cloud resources and Implement Terraform state management and remote backends * Design and implement configuration management solutions using Ansible * Develop Ansible playbooks and roles for automated server provisioning and configuration * Integrate Terraform and Ansible workflows for end-to-end infrastructure automation * Implement infrastructure version control, code review processes, and GitOps practices * Manage infrastructure drift detection and remediation * Create and maintain infrastructure documentation and architecture diagrams * Implement policy-as-code using tools like OPA (Open Policy Agent) or Sentinel CI/CD Pipeline Management * Design, implement, and maintain continuous integration pipelines using Jenkins and Harness * Optimize build times and pipeline efficiency * Integrate security scanning (SAST, DAST, container scanning) into CI/CD pipelines * Configure Jenkins jobs, pipelines, and shared libraries for automated build, configure build agents, runners, and execution environments * Implement Harness deployment pipelines for cloud-native applications * Integrate CI/CD pipelines with version control systems (Git, GitHub, GitLab) * Implement continuous deployment workflows using ArgoCD for Kubernetes-based applications * Design and implement GitOps workflows with ArgoCD for declarative application delivery * Manage ArgoCD application definitions, sync policies and multi-cluster deployments * Implement progressive delivery strategies (blue-green deployments, canary releases) using ArgoCD Message Streaming & Event-Driven Architecture * Deploy and manage Apache Kafka clusters for real-time data streaming and event-driven architectures * Configure Kafka topics, partitions, replication factors, and retention policies * Implement Kafka Connect for data integration with various sources and sinks * Monitor Kafka cluster health, performance metrics, and consumer lag * Optimize Kafka performance for high-throughput and low-latency use cases * Troubleshoot Kafka producer and consumer issues Database & Proxy Management * Deploy, configure, and manage ProxySQL for MySQL load balancing and high availability * Implement query routing, caching, and connection pooling strategies using ProxySQL * Optimize database performance through ProxySQL query analysis and optimization * Implement database failover and disaster recovery using ProxySQL * Monitor ProxySQL metrics and troubleshoot connection and performance issues * Integrate ProxySQL with database clusters and replication topologies * Implement database access security and audit logging through ProxySQL Cloud Networking * Design and implement cloud networking architectures, including VPCs, subnets, and network segmentation * Configure and manage cloud load balancers (Application Load Balancers, Network Load Balancers, Cloud Load Balancing) * Implement VPN connections, Direct Connect/Interconnect, and hybrid cloud networking solutions * Implement network security controls, including security groups, network ACLs, and firewall rules * Implement network monitoring and traffic analysis * Troubleshoot complex networking issues across multi-cloud environments * Design and implement private connectivity between cloud providers Secrets Management & Security * Configure and manage HashiCorp Vault for centralized secrets management across multi-cloud environments * Configure Vault secret engines (KV, database, PKI, AWS, GCP, Azure dynamic secrets) * Manage Vault high availability clusters and disaster recovery procedures * Implement dynamic database credentials and secret rotation strategies * Manage Vault encryption as a service for application-level encryption * Implement Vault agent and sidecar injectors for Kubernetes workloads * Migrate secrets from legacy systems to Vault Qualifications, Competency & Skills Required Education & Experience * Bachelor's degree or diploma in Computer Science, Information Technology, Engineering, or related field * Minimum of 5 years of proven experience in cloud engineering, DevOps, or platform engineering roles * Hands-on experience managing production workloads across multiple cloud platforms * Relevant cloud and technology certifications are highly desirable Technical Skills Cloud Platforms (Required) * Google Cloud Platform (GCP): Deep expertise in Compute Engine, GKE, Cloud Storage, Cloud SQL, VPC, Cloud Functions, Cloud Run, IAM * Amazon Web Services (AWS): Proficiency in EC2, EKS, S3, RDS, VPC, Lambda, ECS, CloudFormation, IAM * Microsoft Azure: Experience with Virtual Machines, AKS, Blob Storage, Azure SQL, Virtual Networks, Azure Functions, ARM templates * Oracle Cloud Infrastructure (OCI): Familiarity with Compute, OKE, Object Storage, networking, and OCI-specific services * Multi-cloud architecture design and implementation experience * Cloud migration strategies and execution (lift-and-shift, re-platforming, re-architecting) Container & Orchestration (Required) * Expert-level Kubernetes knowledge, including cluster architecture, networking, storage, and security * Hands-on experience with managed Kubernetes services (GKE, EKS, AKS) * Proficiency in Docker containerization, image optimization, and registry management * Experience with Helm charts for application packaging and deployment * Knowledge of container runtime environments (containerd, CRI-O) Service Mesh & Microservices (Required) * Istio: Deep expertise in Istio architecture, deployment, and operations * Istio traffic management (virtual services, destination rules, gateways, service entries) * Istio security features (mTLS, authorization policies, peer authentication, request authentication) * Istio observability and telemetry configuration * Multi-cluster and multi-mesh deployments * Service mesh troubleshooting and performance optimization * Understanding of sidecar proxy patterns and Envoy proxy * Experience with other service mesh solutions (Linkerd, Consul Connect) is a plus Reverse Proxy & Load Balancing (Required) * HAProxy Advanced configuration and management for load balancing and high availability * Nginx Expert-level configuration as reverse proxy and API gateway * Nginx rate limiting, and performance tuning * Nginx load balancing algorithms and upstream configurations * Experience with Nginx modules and custom configurations * High availability configurations using keepalived, VRRP, or similar * Integration with Kubernetes ingress controllers (Nginx Ingress, Istio Ingress) Infrastructure as Code (Required) * Advanced Terraform skills for multi-cloud infrastructure provisioning * Terraform module development, state management, and workspace strategies * Proficiency in Ansible for configuration management and automation * Ansible playbook development, roles, and inventory management * Experience with version control systems (Git) and GitOps workflows * Infrastructure testing frameworks (Terratest, Kitchen-Terraform) CI/CD Tools (Required) * Jenkins: Pipeline development (declarative and scripted), shared libraries, plugin management * Harness: Deployment pipeline configuration, workflow creation, approval gates * ArgoCD: GitOps workflows, application synchronization, multi-cluster management * Integration of CI/CD tools with Kubernetes and cloud platforms * Automated testing and deployment strategies * Artifact repository management (Nexus, Artifactory, cloud-native registries) Messaging & Streaming (Required) * Apache Kafka architecture, cluster management, and operations * Kafka topic design, partitioning strategies, and performance tuning * Kafka Connect experience * Experience with Kafka management tools (Kafka Manager, Cruise Control) * Understanding of event-driven architectures and patterns Database & Proxy Technologies (Required) * ProxySQL configuration, management, and optimization * MySQL database administration basics * Understanding of database replication and clustering Observability & Monitoring (Required) * Prometheus metrics collection, PromQL, and alerting rules * Grafana dashboard design and visualization techniques * Log aggregation and analysis Networking (Required) * Deep understanding of TCP/IP, DNS, HTTP/HTTPS, and network protocols * Cloud networking concepts (VPC, subnets, routing tables, NAT, VPN) * Load balancing strategies and implementations * Service discovery and DNS-based routing * Network security and firewall configuration * Software-defined networking (SDN) concepts Scripting & Programming * Proficient in scripting languages: Python, Bash * Go or python programming basics for tooling development * YAML and JSON for configuration management * Understanding of software development best practices Secrets Management (Required) * HashiCorp Vault: Advanced knowledge of Vault architecture, deployment, and operations * Vault authentication methods and integration with cloud providers and Kubernetes * Vault secret engines (KV v1/v2, database, transit, cloud dynamic... What we can offer you * Culture -We put our people first and prioritize the well-being of every team member. We’ve built a company where all opinions carry weight and where all voices are heard. We value and respect each other and always look out for one another. Above all, we are human. * Learning - We have a learning and development-focused environment with an emphasis on knowledge sharing, training, and regular internal technical talks. * Compensation - You’ll receive an attractive salary, pension, health insurance, annual bonus, plus other benefits. What to expect in the hiring process * A technical interview with the Hiring Manager * A behavioural and technical interview with a member of the Executive team.
Camunda is the enterprise platform for agentic orchestration, enabling organizations to coordinate AI agents, people, and systems across complex, end-to-end business processes. With built-in governance, auditability, and human oversight, Camunda gives enterprises the control they need to move AI from pilots to production — safely and at scale. Trusted by over 700 organizations worldwide, including 9 of top 10 US banks, Camunda helps enterprises boost operational efficiency, accelerate time-to-value, and deliver better customer experiences. Fully remote and global, we are in the middle of something bigger: transforming into an AI-first organisation, built on our own platform. We use Agentic AI to automate, orchestrate intelligent processes, and elevate human contribution across every team. Named GP Bullhound’s Top 100 Next Unicorn list, 2025 Great Place to Work certified. Visionary in 2025 Gartner® Magic Quadrant™ for Business Orchestration and Automation Technologies. ranked 3rd in Flexa's 2026 Most Flexible Companies, We’re growing fast and looking for top talent to join our team. If you want meaningful work, visible impact and put something genuinely rare on your CV, keep reading. About the Role: At Camunda, we build the platform behind mission-critical process orchestration for customers worldwide. This role is central to that work. The Manager, Cloud Infrastructure Engineering leads the team that runs the cloud infrastructure behind Camunda's SaaS offering, keeping it reliable, scalable, and secure as the company grows. This is a hands-on engineering leadership role for someone who has built and run Kubernetes-based, multi-cloud infrastructure in production and can help a team do the same. Working fully remote across time zones, this person creates clarity and keeps the team moving. We want someone who uses AI regularly in their infrastructure work, for automation, faster analysis, and tighter operations, and who has thought carefully about where guardrails belong. What you'll be doing: * Lead and grow the team that runs Camunda's SaaS cloud infrastructure, with a focus on execution and strong engineering culture. * Guide the design, delivery, and evolution of Kubernetes-based, multi-cloud infrastructure, with reliability, scalability, and security as the baseline. * Work with product engineering, security, support, and other engineering leaders to make sure the platform lets teams ship safely and efficiently. * Keep raising the bar on infrastructure automation, developer tooling, and platform guardrails so the service scales without adding friction. * Define how the team uses AI as a repeatable part of infrastructure work: ops analysis, automation, incident response, documentation. * Balance day-to-day operational stability with longer-term investments in multi-cloud expansion and new infrastructure capabilities. What you bring: * Ability and/or willingness to use our product. * 3 - 5 years experience leading engineers who build and run production infrastructure for a SaaS product. * Strong technical background in cloud infrastructure engineering, platform engineering, SRE, or a related domain. * Hands-on Kubernetes experience in production, including scaling, reliability, and cluster lifecycle. * Experience working across more than one major cloud provider and understanding the trade-offs of running multi-cloud. * Track record of working with senior engineers and stakeholders across teams to deliver major infrastructure improvements while raising reliability, scalability and keeping a focus on security. * Strong people leadership: coaching, managing performance, setting priorities, and building an environment where engineers can do their best work. * Clear thinking about where AI fits in infrastructure engineering work, with judgment about what's useful, repeatable, and responsible. Nice-to-haves: * Experience running global or multi-region SaaS infrastructure. * Familiarity with infrastructure as code, policy as code, and platform automation. * Experience improving observability, incident management, or reliability for cloud-native systems. * Experience with cloud cost management, platform guardrails, or other tools that help teams make better infrastructure decisions. This role is an existing vacancy #LI-SK1 #LI-Remote #C1 What We Have to Offer: Compensation We offer competitive, fair, and transparent compensation. Salary ranges are location-based, with Standard and Major markets (global tech hubs) reflecting local competition. The Annual Total Target Cash (base salary + 100% variable target, where applicable) shown below spans from the minimum in a Standard market to the maximum in a Major market. Final offers depend on skills, experience, and location, and we typically hire in the first half of the range to allow room for growth: United States: $172,600.00 to $278,300.00 United Kingdom: £108,400.00 to £178,300.00 Singapore: S$214,400.00 to S$321,500.00 If you’re based elsewhere, you’ll be hired via Remote.com (our global employer partner), and your Talent Acquisition Partner will provide a personalized Total Rewards Calculator after your first interview. Equity: We also offer equity (where applicable) through our Virtual Stock Option Plan (VSOP). Benefits & Perks We invest in your wellbeing, growth, and ability to connect, along with perks that support you no matter where you’re based. Our benefits are globally designed and locally delivered where applicable. * Remote & Flexible: Work from anywhere with the setup that suits you, home office budget, co-working space support, and flexible time off to recharge when you need it. * In Person Connection: We invest in meaningful face time through our Annual Kickoff (Vienna in 2025, Madrid in 2026!), team offsites, and Camundi Connection Budgets, including contributing to meetups while travelling,, and local gatherings with fellow Camundi. * Health & Wellbeing: Access locally tailored healthcare, Modern Health for global mental wellbeing, and our Live Well Lifestyle Spending Account (LSA), a flexible, global benefit that puts you in control of your whole life, not just work, from: staying active, to caring for family, exploring personal passions, meaningful experiences, and investing in your financial wellbeing. The Live Well program launches in 2026 and scales to €1,000 annually from 2027. * Financial Security: Retirement and pension plans (often with company contributions), plus life and disability insurance where relevant. * Professional Growth: Up to $/€/£1,000 per year for self-driven learning: courses, certifications, books, you decide! ”Everyone is welcome at Camunda” — it’s a celebrated component of our culture. We strive to create an inclusive environment that empowers our people. At Camunda, we honour diverse cultures and backgrounds and are proud to be an equal opportunity employer. All qualified applicants will receive consideration without regard to gender, race, ethnicity, religion, belief, sexual orientation, age, disability or any other protected characteristics under applicable law. We are looking forward to your application! Come join us and be part of Camunda’s incredible journey: Make an impact at a pivotal moment in our story! AI in our hiring process: Camunda may use AI tools to aid the screening of applications and during the interview process. You can learn more here