
Veritaz AB · Sweden
Experienced Major Incident Manager wanted to ensure rapid escalation and coordinate technical teams during business-critical IT incidents in Solna.
Experienced Major Incident Manager wanted to ensure rapid escalation and coordinate technical teams during business-critical IT incidents in Solna.
Veritaz is a leading IT staffing solutions provider in Sweden, committed to advancing individual careers and aiding employers in ensuring the perfect talent fit. With a proven track record of successful partnerships with top companies, we have rapidly grown our presence in the USA, Europe, and Sweden as a dependable and trusted resource within the IT industry.
Assignment Description
We are looking for an experienced Major Incident Manager
What You Will Work On
Lead and coordinate the Major Incident Management process from detection to resolution
Manage high-priority and business-critical IT incidents
Ensure rapid escalation to appropriate technical teams and stakeholders
Coordinate cross-functional response teams during major incidents
Drive incident communications to operational teams, management, and business stakeholders
Monitor incident progress and ensure timely resolution
Facilitate decision-making during critical situations
Conduct post-incident reviews and root cause analysis (RCA)
Identify and drive corrective and preventive actions
Ensure adherence to ITIL processes and organizational governance models
Coordinate and manage external suppliers during incident resolution activities
Monitor supplier performance against SLAs and service commitments
Support operational readiness and incident response capabilities
Participate in on-call/standby rotations as agreed
Contribute to continuous improvement initiatives within Incident Management and IT Operations
What You Bring
Proven experience in Incident Management or Major Incident Management within large-scale IT environments
Strong understanding of IT Service Management (ITSM) principles and frameworks
Experience managing business-critical incidents with significant operational impact
Experience coordinating technical teams without direct line-management responsibility
Strong stakeholder management and communication skills
Experience working in multi-vendor and outsourced service environments
Ability to remain calm and decisive under pressure
Experience conducting Root Cause Analysis (RCA) and driving service improvements
Strong understanding of complex IT infrastructures, applications, and service dependencies
Ability to communicate effectively with both technical teams and executive stakeholders
Fluent Swedish and English, written and spoken
TITLE: MAJOR INCIDENT MANAGER STATUS: EXEMPT REPORTS TO: MGR – IT SERVICE MANAGEMENT DEPARTMENT: IT – SERVICE MANAGEMENT JOB CODE: 10086 PAY RANGE: $102,000.00 - $125,000.00 ANNUALLY GENERAL DESCRIPTION: As a Major Incident Manager within the IT Operations department, you will manage our Major Incident Management Process across the entire organization. This includes facilitation of technical and application recovery bridges, documenting recovery timelines, executing playbooks for outages, ensuring teams follow incident protocols, and running point on incident communication plans. The Major Incident Manager will also manage the problem management process. This includes managing post major incident activities of root cause analysis, resulting action items, and updating documentation. TASKS, DUTIES, FUNCTIONS: 1. Ensuring flawless execution of the incident resolution process, with transparent communication that drives very high levels of internal/external customer satisfaction. 2. Creating, communicating, and executing the incident response strategy and actions for individual incidents. 3. Managing resources assigned to the incident and ensures the incident is receiving the proper support to drive resolution as quickly as possible. 4. Escalating, prioritizing, communicating, and coordinating high severity incidents ensuring adherence to the company’s incident response process. 5. Addressing incoming escalations from executives regarding the incident. 6. Ensuring all agreed to operational policies and procedures are adhered to and championing the incident response process in alignment with ITIL processes. 7. Driving the incident response process from detection through containment and eradication. 8. Leading the coordination with internal stakeholders through resolution of the incident and root cause analysis. 9. Closely partnering and collaborating with Infrastructure, Engineering, Operations, Technical Support, and Leadership to ensure alignment across the business. 10. Leading cross-functional post-incident process reviews to ensure continuous improvement of operations and execution. 11. Contribute to the improvement of the incident response process based on lessons learned. 12. Train and mentor staff on the incident response process. 13. Ensures that the Major Incident Management process is documented and updated. 14. Remains updated on the latest industry standards in the Major Incident Management process area. 15. Submits change requests to change management as required to eliminate know problems. 16. Recording, managing, and advancing the problem by escalating to the elevated level of expertise, if appropriate, by integrating with change management, incident management, and configuration management. 17. Creating tasks for other IT departments to work on the problem resolution 18. Monitors problem resolution and closure. 19. Analyses historical data to identify and eliminate potential incidents before they occur 16. Develops metrics and reporting requirements and maintains relevant SLA/KPI metrics for both the incident and problem processes. 17. Serve as backup and assistance for other ITSM process owners supporting change, asset and knowledge management. 18. Maintains a thorough understanding of state and federal laws and regulations related to credit union compliance including bank secrecy and anti-money laundering laws appropriate to the position. 19. Prior experience in a 24x7x365 operations environment. Experience with Mass notification tools desired 20. Performs other job-related duties as necessary. PHYSICAL SKILLS, ABILITIES, AND EXERTION UTILIZED IN THE PERFORMANCE OF THESE TASKS: 1. Effective oral and written communication skills required to train, direct, and evaluate staff, diagnose, correct and log computer projects and problems. 2. Strong organizational and time management skills; ability to articulate system methodologies and concepts; communicate effectively in providing technical guidance and expertise to other staff 3. Must possess sufficient manual dexterity to skillfully operate applicable computer hardware, a variety of hand tools and standard office equipment. ORGANIZATIONAL CONTACTS & RELATIONSHIPS: 1. INTERNAL: All levels of staff and management 2. EXTERNAL: Vendors QUALIFICATIONS: 1. EDUCATION: Bachelor’s degree in Computer Science, Management Information Systems or comparable discipline or equivalent work experience. 2. EXPERIENCE: * Bachelor’s degree in IT or related professional experience in field * 5+ years’ experience in the Information Security field, including operational security monitoring or incident response experience. * 3+ years managing, coordinating, and ensuring resolution of security issues. * Deep experience leading and responding to complex critical incidents security, availability, or customer experience incidents). * Broad information security knowledge, including some familiarity with key regulations and standards relating to security incident response (e.g., PCI-DSS, GDPR, ISO 27001). * Ability to manage and constantly triage multiple security incidents, differentiating urgent issues from the merely important. * Ability to stand back from a complex problem, logically assess the facts, and formulate a plan of action - even in the worst of situations. * Strong operational and services experience in a cloud services delivery environment. * Strong technical knowledge of complex systems, ideally in a cloud environment. * Strong technical understanding of network fundamentals and common Internet protocols. * Strong technical understanding of the information security threat landscape (attack vectors and tools, best practices for securing systems and networks, etc.). * Superior verbal and written communication skills, including the ability to effectively and clearly communicate complex scenarios to non-technical colleagues. * Strong ability to prioritize and complete multiple tasks during time critical and high-pressure situations * Excellent time management, communication, and emotional intelligence skills * Strong analytical skills to support the synthesis of knowledge and information out of abstract and varying types of data points PHYSICAL REQUIREMENTS: 1. Prolonged sitting throughout the workday to accomplish tasks. 2. Occasional travel may be required. 3. Lift and carry communications equipment and computer hardware weighing up to fifty pounds. 4. Corrected vision in the normal range required to configure, test, and troubleshoot network server hardware and data. 5. Hearing within normal range. 6. Must possess sufficient manual dexterity to skillfully operate applicable computer hardware, a variety of hand tools and standard office equipment. 7. May work additional work hours to accomplish tasks. LICENSES/CERTIFICATIONS: 1. ITIL Foundation mandatory. THIS JOB DESCRIPTION IN NO WAY STATES OR IMPLIES THAT THESE ARE THE ONLY DUTIES TO BE PERFORMED BY THIS EMPLOYEE. HE OR SHE WILL BE REQUIRED TO FOLLOW OTHER INSTRUCTIONS AND TO PERFORM OTHER DUTIES REQUESTED BY HIS OR HER SUPERVISOR THAT ARE WITHIN HIS / HER KNOWLEDGE, SKILL AND ABILITY AS WELL AS HIS / HER MENTAL AND PHYSICAL ABILITIES. #LI-Hybrid
WHAT WE NEED The Operations Lead is responsible for leading the day-to-day operations, service delivery, reliability, and continuous improvement of the Sakani and Ejar platforms. The role ensures platform availability, operational excellence, incident management, disaster recovery readiness, and effective collaboration between business, development, infrastructure, security, QA, and external vendors WHAT YOU'LL DO Operations Management * Lead and manage the operational support of Sakani and Ejar platforms. * Ensure platform availability, stability, performance, and reliability. * Oversee production and non-production environments. * Monitor operational KPIs, SLAs, and service health metrics. * Drive continuous service improvement initiatives. Incident & Problem Management * Lead Major Incident Management activities and coordinate resolution efforts. * Manage incident escalations across Development, Infrastructure, Database, Security, and Vendor teams. * Conduct Root Cause Analysis (RCA) and ensure implementation of corrective actions. * Track recurring issues and drive long-term solutions. Release & Change Management * Oversee production deployments, releases, and maintenance activities. * Ensure operational readiness before major releases. * Review implementation and rollback plans. * Ensure compliance with change management processes and governance standards. Platform & DevOps Operations * Collaborate with DevOps teams to enhance automation, reliability, and operational efficiency. * Ensure monitoring, logging, alerting, and observability capabilities are maintained. * Support Kubernetes, cloud infrastructure, CI/CD pipelines, and platform operations. * Drive operational excellence through automation and process optimization. Disaster Recovery & Business Continuity * Lead Disaster Recovery (DR) planning, testing, and execution activities. * Ensure Recovery Time Objective (RTO) and Recovery Point Objective (RPO) targets are achieved. * Maintain operational runbooks and recovery procedures. * Coordinate periodic DR drills and readiness assessments. Stakeholder Management * Act as the primary operational contact for business and technical stakeholders. * Coordinate with internal teams, external vendors, and service providers. * Prepare and present operational reports, service reviews, and executive updates. * Facilitate operational governance and service review meetings. Vendor & Service Management * Manage vendor performance against agreed Service Level Agreements (SLAs). * Coordinate operational activities with third-party providers. * Handle escalations and ensure timely issue resolution. Performance & Capacity Management * Monitor platform performance, utilization, and capacity trends. * Identify and address performance bottlenecks. * Plan future capacity requirements and scalability improvements. Security, Risk & Compliance * Ensure compliance with organizational policies, security standards, and regulatory requirements. * Support security audits, risk assessments, and compliance initiatives. * Track and mitigate operational risks. Key Performance Indicators (KPIs) * Platform Availability ≥ 99.9%. * SLA Compliance ≥ 95%. * Change Success Rate ≥ 98%. * Reduction in Mean Time to Recovery (MTTR). * Successful Disaster Recovery testing and execution. * Operational Risk Reduction. * Stakeholder and Customer Satisfaction. Reporting To IT Operations Manager / Head of Technology Operations. Scope Responsible for operational governance, service reliability, incident management, platform operations, disaster recovery readiness, and service delivery across. WHAT YOU HAVE Qualifications * Bachelor’s degree in Computer Science, Information Technology, Engineering, or a related field. * Minimum 8 years of experience in IT Operations, Service Delivery, Infrastructure, or DevOps. * Minimum 3 years of experience in a leadership or management role. * Proven experience managing large-scale enterprise platforms and critical business services. Experience working with cross-functional teams and external vendors.
INCIDENT MANAGER US REMOTE CSQ127R151 At Databricks, we are passionate about empowering data teams to tackle the world’s most challenging problems — from bringing the next mode of transportation to reality to accelerating the development of medical breakthroughs. We achieve this by building and operating the world’s best data and AI infrastructure platform, enabling our customers to leverage deep data insights and enhance their business. Founded by engineers — and customer-obsessed — we leap at every opportunity to tackle technical challenges, from designing next-gen UI/UX for interfacing with data to scaling our services and infrastructure across millions of virtual machines. And we're only getting started. As an Incident Manager, you will lead Databricks’ most critical production incidents while providing clear, accurate, and timely communication to customers, executives, and engineers. You’ll serve as both incident commander and reliability engineer; orchestrating multi-team responses, driving real-time status updates, and partnering with engineering to analyze and prevent failures. Your work will ensure Databricks maintains not only technical resilience but also customer and stakeholder confidence during high-impact events. This role combines operational leadership, technical systems knowledge, and exceptional communication skills. You will be at the intersection of engineering depth and operational clarity, ensuring that every major incident is managed with precision, transparency, and continuous improvement. THE IMPACT YOU WILL HAVE HERE: * Lead critical incidents — coordinate multi-disciplinary response efforts across Databricks’ cloud-based services to rapidly mitigate impact and restore operations. * Drive technical root cause analysis and reliability improvements: * collaborate with engineering teams to trace and document underlying causes across distributed systems, services, and data stores. * Summarize key learnings, clearly communicate action items, and ensure that technical and procedural improvements are followed through. * Own communications during incidents — deliver frequent, high-quality updates to internal stakeholders (executives, engineering leadership, support) and compose and publish customer-facing notifications that are accurate, timely, and empathetic. * Mentor and train peers in both incident communication and technical response disciplines to raise the overall quality of Databricks’ incident response. WHAT ARE WE LOOKING FOR: * 5+ years of experience in incident management, site reliability engineering, or production operations supporting large-scale, cloud-native systems. * Proven ability to lead and coordinate high-severity incidents, including identifying impact, isolating fault domains, and managing multi-team response efforts. * Strong understanding of cloud infrastructure (AWS, Azure, or GCP) — including compute, networking, storage, and observability components. * Deep expertise in log analysis and debugging: * Familiarity with log aggregation and search tools (e.g., Datadog, Elasticsearch, Splunk, Cloud Logging, or OpenTelemetry). * Hands-on experience with observability systems — metrics, logging, and tracing frameworks (Prometheus, Grafana, OpenTelemetry, etc.). * Proficiency in at least one major programming or scripting language (Python, Go, or Bash) for automating diagnostics, data collection, or analysis. * Experience developing and maintaining incident playbooks and communication templates to ensure consistent, timely updates. * Excellent contextual interpretation and writing skills, as well as the ability to effectively summarize and communicate to both technical and business audiences, are required. * BS, Master's or other advanced degree in Computer Science or Computer Engineering, or related Engineering field. Pay Range Transparency Databricks is committed to fair and equitable compensation practices. The pay range(s) for this role is listed below and represents the expected salary range for non-commissionable roles or on-target earnings for commissionable roles. Actual compensation packages are based on several factors that are unique to each candidate, including but not limited to job-related skills, depth of experience, relevant certifications and training, and specific work location. Based on the factors above, Databricks anticipates utilizing the full width of the range. The total compensation package for this position may also include eligibility for annual performance bonus, equity, and the benefits listed above. For more information regarding which range your location is in visit our page here. Zone 3 Pay Range $103,900—$145,525 USD About Databricks Databricks is the data and AI company. More than 10,000 organizations worldwide — including Comcast, Condé Nast, Grammarly, and over 50% of the Fortune 500 — rely on the Databricks Data Intelligence Platform to unify and democratize data, analytics and AI. Databricks is headquartered in San Francisco, with offices around the globe and was founded by the original creators of Lakehouse, Apache Spark™, Delta Lake and MLflow. To learn more, follow Databricks on Twitter, LinkedIn and Facebook. Benefits At Databricks, we strive to provide comprehensive benefits and perks that meet the needs of all of our employees. For specific details on the benefits offered in your region click here. Our Commitment to Diversity and Inclusion At Databricks, we are committed to fostering a diverse and inclusive culture where everyone can excel. We take great care to ensure that our hiring practices are inclusive and meet equal employment opportunity standards. Individuals looking for employment at Databricks are considered without regard to age, color, disability, ethnicity, family or marital status, gender identity or expression, language, national origin, physical and mental ability, political affiliation, race, religion, sexual orientation, socio-economic status, veteran status, and other protected characteristics. Compliance If access to export-controlled technology or source code is required for performance of job duties, it is within Employer's discretion whether to apply for a U.S. government license for such positions, and Employer may decline to proceed with an applicant on this basis alone.