
Graphcore · Austin
We are looking for a disciplined and dynamic, Lead System Engineer – compute blade and rack Validation to join our growing compute rack validation team. As a di...
We are looking for a disciplined and dynamic, Lead System Engineer – compute blade and rack Validation to join our growing compute
rack validation team. As a diligent leader in Systems Engineering, you will drive multiple aspects of validation throughout the
life cycle of the program. In this high visibility position, you will be part of a leading team to innovate and improve system
bring-up and enablement abilities, as well as silicon and system validation to deliver the highest quality, industry leading
technologies to market. Your technical leadership skills, validation and debug expertise will be necessary towards product
development, definition, root cause and resolution. Your agility and collaborative approach will be essential to work within
System Validation & other engineering teams (System Architects, SoC and Rack FW etc).
The technical leader will be driving keys areas of system validation including leading first silicon & system bring-up (nodes and
rack level systems) - rack level systems and blades will be based of ARM server architecture. Candidate will be immersed in
challenging system enablement work, system validation (end-to-end) methodology, tests development and execution as well as
triage/debug of critical issues to meet critical program milestones at POR quality. The candidate will also be a key contributor
to state-of-the-art HW and lab capabilities for Grapchore’s system engineering. The candidate should be able to work in a global
environment while maintaining a synergetic culture.
plan of record and system architecture spec.
release of deployment ready solutions.
successful system (HW/SW/FW) bring-up and system validation at blade and rack level for AI compute rack.
Ensure issues are solved on time with quality.
memory, HBM, IO etc.
procedural methodology enhancement, and various internal and cross-functional technical initiatives.
DHCP etc) in a lab environment to help add end-to-end validation and debug capabilities for rack and blade validation.
revision control systems.
interactions.
schedule, quality, or coverage.
certification processes of these environments
challenges in a server compute rack or data center blade environment.
USA Benefits
In addition to a competitive salary, Graphcore offers flexible working and a comprehensive benefits package designed to support
your health, wellbeing and financial future. Our benefits include medical, dental and vision coverage, Flexible Spending Accounts
(FSAs), Health Savings Accounts (HSAs), disability and life insurance, a 401(k) retirement plan, commuter benefits, wellness
services and an Employee Assistance Programme (EAP). We welcome people of different backgrounds and experiences; we're committed
to building an inclusive work environment that makes Graphcore a great home for everyone. We offer an equal opportunity process
and understand that there are visible and invisible differences in all of us. We can provide a flexible approach to interview and
encourage you to chat to us if you require any reasonable adjustments.
ABOUT US We are looking for a disciplined and dynamic Systems Engineer with focus on server CPU based system to join our growing compute rack validation team. Candidate we are seeking should have demonstrated work-experience in leading server rack and blade hardware systems deployment, hardware installation, and inventory management activities in the Austin, TX area. As a diligent leader in Systems Engineering, you will drive multiple aspects of post-silicon validation throughout the life cycle of the program. In this high visibility position, you will be part of a technical team chartered to innovate and improve system bring-up and enablement capabilities, as well as silicon and system validation to deliver the highest quality, industry leading technologies to market. Your technical leadership skills, systems engineering and hardware bring-up, validation and debug expertise will be necessary towards product development, definition, root cause and resolution. Your agility and collaborative approach will be essential to work within System Validation & other engineering teams (System Architects, SoC and Rack FW etc). The ideal candidate will be driving key areas around at-scale system validation including ARM based server and rack level systems bring-up (nodes and rack level systems). Candidate will be immersed in challenging system enablement work, ramp-up post-silicon capabilities in engineering lab environments, validation tests execution/triage. The candidate will be leading contributor towards state-of-the-art HW bring-up and lab capabilities for Grapchore’s system engineering. The candidate should be able to work in a global environment while maintaining a synergetic culture. Primary Responsibilities: * Install, configure, commission (and decommission if needed) blade servers, chassis, switches, and supporting infrastructure. * Lead rack and stack activities, including mounting equipment, cable management, and labelling. Execute hardware upgrades, replacements, and troubleshooting of server and network components. * Maintain accurate asset records within DCIM platforms and inventory management systems. * Conduct physical audits and reconcile inventory discrepancies. * Track hardware movements, deployments, and decommissions through established change management processes. * Document installation procedures, rack layouts, cabling diagrams, and inventory updates. * Support data center migration, expansion, and refresh projects. * Collaborate with engineering, operations, logistics, and project management teams. * Adhere to all data center safety, security, and operational standards. * Develop, setup and scale key methodologies for at-scale test execution, lab HW and system SW capabilities as well as system visibilities and debug tools necessary for successful system (HW/SW/FW) bring-up and system validation at blade and rack level for AI compute rack. * Ability to work independently in a production ready environment, and a commitment tomaintaining accurate inventory and asset records. * Triage issues found during server rack validation bring-up, Post-Silicon Validation, and production phases of the program. Ensure issues are solved on time with quality. * Lead test execution of key domains within AI compute solutions like CPU, GPU, memory, HBM, IO etc. * Drive technical innovation to improve capabilities across system validation, including tools, script development, technical and procedural methodology enhancement, and various internal and cross-functional technical initiatives. Qualifications: * Strong analytical/problem-solving skills and pronounced attention to details * Experience in Blade server installation and maintenance (Cisco UCS, HPE Synergy, Dell MX, or similar). * Rack and stack deployments in enterprise or hyperscale environments. * Copper and fiber cabling installation and management. * DCIM and asset management platforms. * Strong understanding of server, storage, and networking hardware. * Experience performing inventory audits and maintaining asset accuracy. * Ability to read rack elevation diagrams, cabling schematics, and deployment documentation. * Familiarity with ticketing and change management systems. * Exposure to Linux (ubuntu) OS bootable images and system firmware basics for image building, provisioning and firmware flashing. * Exposure to automation testing, to enable execution of hardware acceptance tests, best-known-config testing etc. * Exposure to python script development and execution. * Proven experience in understanding, defining and enabling storage (storage rack), networking capabilities (network rack, DNS, DHCP etc) in a lab environment to help add end-to-end validation and debug capabilities for rack and blade validation. * Excellent communication and coordination skills. * Detailed oriented, highly organized, able to prioritize, and juggle multiple work streams to tight deadlines. * Technical leadership: capable of championing new tools, methods, and capabilities to drive platform validation improvements in schedule, quality, or coverage. * Experience working with data center technical staff, 3rd party vendors, ODMs etc throughout the life cycle of server system product development. * Must be a self-starter, and able to independently drive tasks to completion Preferred Qualifications: * Masters or PhD in Electrical Engineering, Computer Engineering or a related field. * 10+ years of work experience demonstrating working on complex systems engineering challenges to validate and debug HW-FW-SW challenges in a server compute rack or data center blade environment. * Experience designing and deploying modern AI/ML rack scale systems * Knowledge of industry standards and best practices for hardware development * Familiarity with emerging technologies in AI and Data Center infrastructure. * Comfortable meeting, engaging and collaborating with ODM partners and staffing vendors across the globe. USA Benefits In addition to a competitive salary, Graphcore offers flexible working and a comprehensive benefits package designed to support your health, wellbeing and financial future. Our benefits include medical, dental and vision coverage, Flexible Spending Accounts (FSAs), Health Savings Accounts (HSAs), disability and life insurance, a 401(k) retirement plan, commuter benefits, wellness services and an Employee Assistance Programme (EAP). We welcome people of different backgrounds and experiences; we're committed to building an inclusive work environment that makes Graphcore a great home for everyone. We offer an equal opportunity process and understand that there are visible and invisible differences in all of us. We can provide a flexible approach to interview and encourage you to chat to us if you require any reasonable adjustments.
ABOUT US Graphcore is one of the world’s leading innovators in Artificial Intelligence compute. It is developing hardware, software and systems infrastructure that will unlock the next generation of AI breakthroughs and power the widespread adoption of AI solutions across every industry. As part of the SoftBank Group, Graphcore is a member of a best-in-class family of companies responsible for some of the world’s most transformative technologies. Together, they share a bold vision: to enable Artificial Super Intelligence and ensure its benefits are accessible to everyone. Graphcore’s teams are drawn from a diverse group of backgrounds and bring a broad range of skills and perspectives. A melting pot of AI research specialists, silicon designers, software engineers and systems architects, Graphcore enjoys a culture of continuous learning and constant innovation. JOB SUMMARY We are seeking a senior validation lead engineer to lead at-scale rack validation efforts for next-generation AI hyperscale systems. This role focuses on post-silicon system validation across the full lifecycle, ensuring functional, electrical, and thermal performance meets product objectives. You will own end-to-end blade and rack validation including planning, development, execution, and debug while collaborating across firmware, systems, and hardware teams. THE TEAM The Rack Validation team is responsible for ensuring system readiness and quality at scale. The team works cross-functionally with firmware, silicon, and system engineering teams to validate complex AI compute platforms. RESPONSIBILITIES AND DUTIES * Lead post-silicon validation of AI compute blades and racks including test planning, development, and automation. * Drive provisioning and integration of system components (SoC FW, BMC, RMC, OS) for rack-level readiness. * Own execution against program achievements and report validation progress and risks. * Triage test failures, collect debug data, and collaborate on root cause analysis. * Track validation coverage and continuously improve test processes and infrastructure. * Collaborate with ODM/JDM partners on validation and quality. * Mentor engineers and drive engineering excellence. CANDIDATE PROFILE Essential: * Bachelor’s or Master’s degree or equivalent experience in Computer Engineering, Electrical Engineering, Computer Science, or related field. * Proven track record in system, rack, or embedded validation with leadership experience. * Strong experience in large-scale hardware validation environments. * Expertise in CPU/GPU, memory, IO, and firmware validation. * Experience with Linux/server OS and automation using Python/Bash. * Knowledge of IPMI, Redfish, PLDM. * Experience with CI/CD pipelines and hardware interfaces. Desirable: * Experience in hyperscale environments. * Familiarity with OpenBMC and processes for verifying firmware functionality. * Knowledge of firmware security and HIL testing. * Experience with test management tools. USA Benefits In addition to a competitive salary, Graphcore offers flexible working and a comprehensive benefits package designed to support your health, wellbeing and financial future. Our benefits include medical, dental and vision coverage, Flexible Spending Accounts (FSAs), Health Savings Accounts (HSAs), disability and life insurance, a 401(k) retirement plan, commuter benefits, wellness services and an Employee Assistance Programme (EAP). We welcome people of different backgrounds and experiences; we're committed to building an inclusive work environment that makes Graphcore a great home for everyone. We offer an equal opportunity process and understand that there are visible and invisible differences in all of us. We can provide a flexible approach to interview and encourage you to chat to us if you require any reasonable adjustments.
LEAD SYSTEM DEBUG ENGINEER – SERVER & RACK VALIDATION (PRINCIPAL LEVEL AND ABOVE) POSITION OVERVIEW We are seeking a senior technical leader (Principal Engineer level and above) to lead the bring-up, enablement, and hardware debug of server compute systems and rack-level platforms based on Arm® server architecture. The successful candidate will be a key member of the System Validation organization, responsible for driving system-level debug activities and facilitating rapid resolution of complex hardware, firmware, and software issues. This role requires close collaboration with engineering teams across the organization to identify root causes, implement corrective actions, and ensure successful program execution. The ideal candidate will be deeply involved in challenging system debug efforts while developing and executing scalable debug strategies that maximize throughput and ensure Product of Record (POR) quality. In addition, this individual will establish and drive debug methodologies, improve processes, and help create a culture of technical excellence across the organization. We are looking for a disciplined, dynamic, and highly motivated leader who can thrive in a global environment while fostering strong cross-functional collaboration. As a Lead System Debug Engineer within Server and Rack Validation, you will drive balanced, scalable, and automated debug solutions that optimize engineering efficiency and product quality. This highly visible role provides the opportunity to innovate and improve debugging capabilities while delivering industry-leading server technologies to market. Your technical leadership, validation expertise, and problem-solving skills will play a critical role in product development, issue root cause analysis, and resolution. Success in this role requires close collaboration with System Validation, System Architecture, Silicon Engineering, Rack Firmware, and other cross-functional teams. ---------------------------------------------------------------------------------------------------------------------------------- PRIMARY RESPONSIBILITIES * Develop and drive a Debug Center of Excellence, including scalable debug and triage methodologies, processes, and playbooks for server blade and rack-level issues spanning hardware, firmware, and software integration. * Debug issues discovered during server rack bring-up, post-silicon validation, and production phases. * Lead complex debug efforts involving silicon, server systems, firmware, and software to determine root causes and drive effective resolutions. * Ensure issues are resolved with high quality and within program timelines. * Manage and track technical issues, risks, and priorities to remove blockers and achieve key program milestones. * Develop and publish debug program metrics and indicators to identify roadblocks and improve overall debug efficiency. * Communicate program status, risks, and opportunities to customers, stakeholders, and executive leadership. * Drive technical innovation across triage and debug workflows through tool development, scripting, methodology enhancements, and cross-functional engineering initiatives. * Mentor engineers and promote best practices in system validation and debug methodologies. ---------------------------------------------------------------------------------------------------------------------------------- REQUIRED QUALIFICATIONS * Strong analytical and problem-solving skills with exceptional attention to detail. * Extensive experience in validation and debug roles involving operating systems, firmware, silicon, and hardware issues. * Deep understanding of industry-standard server interconnects and software stacks, including PCIe and CXL. * Strong knowledge of Arm® CPU or x86 architectures, SoC design, memory subsystems, RAS (Reliability, Availability, and Serviceability), and power management. * Extensive experience with system architecture, technical debugging, and validation strategies. * Strong understanding of platform-level and system-level debug methodologies, including Operating Systems, Device Drivers, and BIOS interactions. * Excellent communication, collaboration, and cross-functional leadership skills. * Highly organized and detail-oriented, with the ability to manage multiple priorities and deliver results under tight deadlines. * Experience leading technical programs and coordinating cross-functional engineering efforts. * Thorough understanding of data center technologies and associated software stacks. * Self-motivated with the ability to independently drive tasks from problem identification through resolution. ---------------------------------------------------------------------------------------------------------------------------------- PREFERRED QUALIFICATIONS * Master's degree or Ph.D. in Electrical Engineering, Computer Engineering, Computer Science, or a related technical field. * Experience with large-scale server platforms, rack-level systems, and hyperscale data center environments. * Expertise in automation, scripting, and debug tool development. * Experience establishing and scaling debug processes across multiple product generations and engineering organizations. USA Benefits In addition to a competitive salary, Graphcore offers flexible working and a comprehensive benefits package designed to support your health, wellbeing and financial future. Our benefits include medical, dental and vision coverage, Flexible Spending Accounts (FSAs), Health Savings Accounts (HSAs), disability and life insurance, a 401(k) retirement plan, commuter benefits, wellness services and an Employee Assistance Programme (EAP). We welcome people of different backgrounds and experiences; we're committed to building an inclusive work environment that makes Graphcore a great home for everyone. We offer an equal opportunity process and understand that there are visible and invisible differences in all of us. We can provide a flexible approach to interview and encourage you to chat to us if you require any reasonable adjustments.