
Braiins · Braiins
✨ WHAT AWAITS YOU HERE: As a Site Reliability Engineer, you will help us make Braiins Pool and related products faster, safer, and smoother to develop, deploy,...
As a Site Reliability Engineer, you will help us make Braiins Pool and related products faster, safer, and smoother to develop,
deploy, and operate. You will work close to backend, frontend, support, and platform teams to improve deployment pipelines,
Kubernetes-based delivery, observability, incident handling, internal tooling, and developer experience.
This is a hands-on engineering role for someone who enjoys connecting systems, automation, reliability, and practical software
development. The role can lean more toward deployment automation, CI/CD, SRE practices, backend tooling, or support enablement
depending on your strengths.
production.
with the platform team.
documentation, or better tooling.
strong human ownership, review, security, and production quality.
About Us Visa is a world leader in payments technology, facilitating transactions between consumers, merchants, financial institutions and government entities across more than 200 countries and territories, dedicated to uplifting everyone, everywhere by being the best way to pay and be paid. At Visa, you'll have the opportunity to create impact at scale — tackling meaningful challenges, growing your skills and seeing your contributions impact lives around the world. Join Visa and do work that matters – to you, to your community, and to the world. Progress starts with you. Job Description We are seeking a Site Reliability Engineer (SRE) to join Visa Open Banking (VOB), delivering open banking solutions as part of Visa. Visa Open Banking processes billions of events, enabling data-driven decisions for clients and internal stakeholders. As an SRE in the Infrastructure & Tooling area, you will play a key role in ensuring our platforms are reliable, scalable, secure, and easy to use, while empowering engineering teams with self-service tools and automation. This role spans both cloud infrastructure (AWS, Kubernetes, runtime) and engineering productivity platforms (CI/CD, developer portals, observability), focusing on improving developer experience and overall system resilience. Key Responsibilities: Build and operate reliable, scalable, and secure infrastructure platforms across AWS and Kubernetes Develop and maintain self-service tooling and platforms to enable teams to deploy and operate services independently Improve CI/CD pipelines, developer experience, and platform usability Drive observability strategies across logs, metrics, and tracing Automate provisioning, deployment, and scaling using scripting (Python, Go, Bash, etc.) Support engineering teams with best practices in reliability, performance, and cost optimisation Collaborate with stakeholders to ensure platforms align with evolving use cases and requirements Contribute to architecture, standards, and platform designs for large-scale distributed systems Act as a champion for automation, platform engineering, and operational excellence This is a hybrid position. Expectation of days in office will be confirmed by your hiring manager. Qualifications Basic Qualifications: Relevant work experience and a Bachelor's degree Experience working with cloud platforms (AWS or similar) and container orchestration (Kubernetes/EKS) Solid understanding of CI/CD pipelines and modern software delivery practices Experience with infrastructure as code and automation tools Working knowledge of monitoring and observability tools (e.g. metrics, logs, tracing) Programming or scripting skills (e.g. Python, Go, Bash) Understanding of distributed systems, reliability, and scalability principles Strong collaboration and communication skills Preferred Qualifications: Experience building or operating developer platforms (e.g. Backstage, internal portals, self-service tooling) Experience with observability platforms such as Datadog Knowledge of platform engineering and developer experience practices Experience with large-scale, event-driven systems Understanding of cost optimisation and operational efficiency in cloud environments Experience influencing architecture decisions and engineering standards Familiarity with security and compliance requirements in cloud environments Visa is an EEO Employer Qualified applicants will receive consideration for employment without regard to race, color, religion, sex, national origin, sexual orientation, gender identity, disability or protected veteran status. Visa will also consider for employment qualified applicants with criminal histories in a manner consistent with EEOC guidelines and applicable local law.
About BlaBlaCar BlaBlaCar is the world’s leading community-based travel app enabling 27 million members a year to carpool or travel by bus in 21 countries. Our team of 800 employees counts over 50 nationalities and is spread across our 5 global offices, 30% working fully remotely. Your Mission By joining our Foundations department, you will be working alongside talented individuals grouped in small agile teams that each have strong ownership on their piece of these goals. Foundations is composed of seven teams which “provide consistent, easy to use, infrastructures, services, and expertise to support BlaBlaCar’s growth and evolution”. The Site Reliability Engineering team (SRE) is responsible to provide best in class Observability, Alerting and Incident management tools and processes to service teams. As an enabling team, we help BlaBlacar engineers to efficiently improve their service reliability. Empowering developers and bringing them our reliability expertise are at the core of our daily work. Technical stack: Core Infrastructure: Kubernetes, Google Cloud Platform GitOps/Delivery: GitHub, Terraform, Flux, Helm, Jenkins Observability/Incident Management: Datadog, Opentelemetry, Grafana IRM, In house Synthetic Tests platform: Playwright, Qualcium, SauceLabs Languages: Go / Python for Tooling, Typescripts/JS for the testing platform Your responsibilities Support software engineers by creating, maintaining, and improving observability and alerting tools and frameworks. You embrace the use of AI, leveraging agentic to eliminate toil and streamline your daily tasks Own the Service Level Objectives (SLOs) framework, assist in the design and maintenance of indicators (SLI) and objectives to ensure service reliability. Owning the incident management process by defining best practices, standards, and ensuring continuous improvement through post-mortems and chaos engineering. While developers handle incidents within their scope, you could step in as Incident Commander during high-severity incidents, leading coordination efforts . Develop and maintain tools, such as Terraform modules or Go apps, to help automate and enhance reliability across services. Build and promote reporting on operational metrics and incidents to drive distributed and continuous improvement. Your qualifications 1 to 5 years of experience in SRE, DevOps, or Software Engineering roles Working in a multidisciplinary environment will request strong communication skills : you'll need to adapt your communication level to other teams expertise and be able to understand their needs Strong knowledge of observability tools (e.g., Datadog) and understanding of metrics, logging, and tracing. Troubleshooting/oncall experience in production environments, diagnosing and resolving technical issues effectively (experience with Kubernetes is a plus). Full working proficiency in English Fit with our BlaBlaPrinciples Thriving in a collaborative, fast-growing and innovative environment Ability to take ownership, aligned with business priorities and navigating in different contexts Nice to have: Familiarity with incident management platforms (e.g., Grafana IRM) is a bonus Experience working with Service Level Objectives (SLOs) and Service Level Indicators (SLIs) Exposure to programming in Go or a strong interest in learning it. Experience in integrating Opentelemetry Backend services are built using multiple programming languages: while development skills aren't required, familiarity with object-oriented programming and scripting languages is an advantage. Familiarity with web/mobile testing tools or a strong curiosity to understand how software is tested at scale. What we have to offer Hybrid status for this role : 2-3 days at the Office 4 additional weeks on top of legal maternity/paternity leaves 50% healthcare coverage (Alan) Financial support for home office equipment Minimum 25 days holiday per year Local meal plan policy (Swile card) 50% transportation paid (Forfait Mobilité Durable) Free unlimited carpooling & bus rides Personal growth via trainings, mentorship, and internal mobility opportunities Employee Stock ownership plan Regular team building events 1 day off per year to test our product Interested in joining the ride? a 45-min video-call with Maxime, Talent Acquisition Manager, to get to know you, understand your career expectations and answer your questions a 60-min video-call with Damien Bertau, Hiring Manager, to discuss your experience and share more details about the team a 90-min system design interview with 2 team members to discuss about your technical expertise a 45-min video-call with Maxime Fouilleul, Head of Foundations, to get a wider vision of the department and its strategy Our hiring process lasts on average 25-30 days, offers usually come within 48 hours. Please note that one of these interviews will be onsite.
We’re Capital on Tap 👋 💳 Capital on Tap started because small businesses were underserved. Big banks were slow, their products weren't fit for purpose, and small business owners often couldn't access what they needed. We set out to fix that. Today we're a financial platform - not just a credit card company. We offer a best-in-class business credit card, SME-focused spend management platform, a savings product that hit £1 billion in funds within its first year, and a growing suite of tools and financial products that make running a small business easier. 1,000+ employees, £20bn in annual card spend, 200,000+ customers, 17,000+ Trustpilot reviews averaging 4.7 stars, and we're profitable. We’ve done a pretty good job so far, but we’re just getting started! 📍London, Old Street | 🏢 2 Days in Office SRE at Capital On Tap 🌞 At Capital On Tap, we run a hybrid embedded SRE model. We aim to work closely with the teams within Capital On Tap to provide them the best support. Our main objective currently is to gain as much visibility into our platform's health while offering scalable solutions. What You’ll be doing: As a Site Reliability Engineer (SRE) you will help ensure our platforms are fast, reliable, and scalable. You’ll design, build, and monitor systems, prevent issues before they happen. Using SLAs, SLIs, and SLOs, you’ll guide feature launches while maintaining services that everyone can depend on. * Manage and automate Azure, Datadog, NGINX & Cloudflare * Develop and monitor Kubernetes and Serverless resources * Maintain infrastructure code with Terraform & CRDs / Crossplane * Improve systems, processes, and technologies; consult stakeholders to enhance platform performance * Getting involved in new application architecture & design processes * Design solutions to reduce toil, automate repetitive tasks and streamline workflows to reduce manual work and boost team productivity. * Create SLIs and SLOs; increase application visibility * Align with the Product team on SLAs and core service objectives * Collaborate with Platform Engineers for automated solutions and pipelines * Enhance user experience with infrastructure and pipeline optimisation * Support CI/CD tools such as Azure Devops, Octopus Deploy and Flux to streamline software delivery * Lead incident troubleshooting to safeguard customer experience We’re Looking For 🔎 Required skills: * Experience in managing a public cloud (Azure advantageous) * Experience in Azure DevOps, Octopus, Flux or other CI/CD tools * Experience with Linux and Microsoft Systems * Excellent communication skills and ability to collaborate with multiple teams in an agile environment * Proficient in contributing to IaC technologies involving expertise in writing, managing, and optimising infrastructure with tools such as Terraform and Pulumi * Experience working with a cloud monitoring solution (advantageous to have DataDog) * Experience with Kubernetes and Docker * Experience in at least one scripting language (Python, PowerShell, Go) Interview Process 🤝 * First stage: 30-minute intro, CV review, and values with Talent Partner * Second stage: 60 minute “Tech Chat” with Team Manager * Final stage: 75-minute Technical Task + 30 minute Interview with Head of Platform Engineering Diversity & Inclusion 🌈 We welcome, consider and encourage applications from anyone who shares our commitment to inclusivity. Join us in creating a space where authenticity thrives, and everyone can do their best work. Great Work Deserves Great Perks We try not to take ourselves too seriously (all the time) so we make sure our office is decked out with a pool table, arcade machine, beer tap, and a couple of office dogs thrown in for good measure. Check out our benefits: 🏥 Private Healthcare including dental and opticians services through Vitality ✈️ Worldwide travel insurance through Vitality 🎁 Anniversary Rewards (£250, £500, £750, 4-week fully paid sabbatical) 👛 Salary Sacrifice Pension Scheme up to 7% match 🏖️ 28 days holiday (plus bank holidays) 📖 Annual Learning and Wellbeing Budget 👪 Enhanced Parental Leave 🚲 Cycle to Work Scheme 🚂 Season Ticket Loan 💬 6 free therapy sessions per year 🐶 Dog Friendly Offices 🍫 Free drinks and snacks in our offices Check out more of our benefits, values and mission here. Other Info 👍Check out our ‘Top Tips’ for interviewing. ✔️Keep updated on new job opportunities by following us on Linkedin. 📧Email careers@capitalontap.com if you have any questions. Excited to work here? Apply! If you’d like to progress your career within our fast growing, profitable fintech then click apply and we will aim to get back to you within 3 working days (during busy periods this could take up to 5 working days.)