
Databricks · San Francisco
P-1284 ABOUT THIS ROLE As a software engineer for GenAI inference, you will help design, develop, and optimize the inference engine that powers Databricks’ F...
As a software engineer for GenAI inference, you will help design, develop, and optimize the inference engine that powers
Databricks’ Foundation Model API. You’ll work at the intersection of research and production, ensuring our large language model
(LLM) serving systems are fast, scalable, and efficient. Your work will touch the full GenAI inference stack — from kernels and
runtimes to orchestration and memory management.
large-scale LLMs inference
mixture-of-experts) into the engine
workloads
model versioning
overhead
partitioning
Pay Range Transparency
Databricks is committed to fair and equitable compensation practices. The pay range(s) for this role is listed below and
represents the expected salary range for non-commissionable roles or on-target earnings for commissionable roles. Actual
compensation packages are based on several factors that are unique to each candidate, including but not limited to job-related
skills, depth of experience, relevant certifications and training, and specific work location. Based on the factors above,
Databricks anticipates utilizing the full width of the range. The total compensation package for this position may also include
eligibility for annual performance bonus, equity, and the benefits listed above. For more information regarding which range your
location is in visit our page here.
Local Pay Range
About Databricks
Databricks is the data and AI company. More than 10,000 organizations worldwide — including Comcast, Condé Nast, Grammarly, and
over 50% of the Fortune 500 — rely on the Databricks Data Intelligence Platform to unify and democratize data, analytics and AI.
Databricks is headquartered in San Francisco, with offices around the globe and was founded by the original creators of Lakehouse,
Apache Spark™, Delta Lake and MLflow. To learn more, follow Databricks on Twitter, LinkedIn and Facebook.
Benefits
At Databricks, we strive to provide comprehensive benefits and perks that meet the needs of all of our employees. For specific
details on the benefits offered in your region click here.
Our Commitment to Diversity and Inclusion
At Databricks, we are committed to fostering a diverse and inclusive culture where everyone can excel. We take great care to
ensure that our hiring practices are inclusive and meet equal employment opportunity standards. Individuals looking for employment
at Databricks are considered without regard to age, color, disability, ethnicity, family or marital status, gender identity or
expression, language, national origin, physical and mental ability, political affiliation, race, religion, sexual orientation,
socio-economic status, veteran status, and other protected characteristics.
Compliance
If access to export-controlled technology or source code is required for performance of job duties, it is within Employer's
discretion whether to apply for a U.S. government license for such positions, and Employer may decline to proceed with an
applicant on this basis alone.
P-1285 ABOUT THIS ROLE As a staff software engineer for GenAI inference, you will lead the architecture, development, and optimization of the inference engine that powers Databricks Foundation Model API.. You’ll bridge research advances and production demands, ensuring high throughput, low latency, and robust scaling. Your work will encompass the full GenAI inference stack: kernels, runtimes, orchestration, memory, and integration with frameworks and orchestration systems. WHAT YOU WILL DO * Own and drive the architecture, design, and implementation of the inference engine, and collaborate on model-serving stack optimized for large-scale LLMs inference * Partner closely with researchers to bring new model architectures or features (sparsity, activation compression, mixture-of-experts) into the engine * Lead the end-to-end optimization for latency, throughput, memory efficiency, and hardware utilization across GPUs, and accelerators * Define and guide standards to build and maintain instrumentation, profiling, and tracing tooling to uncover bottlenecks and guide optimizations * Architect scalable routing, batching, scheduling, memory management, and dynamic loading mechanisms for inference workloads * Ensure reliability, reproducibility, and fault tolerance in the inference pipelines, including A/B launches, rollback, and model versioning * Collaborate cross-functionally on Integrating with federated, distributed inference infrastructure – orchestrate across nodes, balance load, handle communication overhead * Drive cross-team collaboration: with platform engineers, cloud infrastructure, and security/compliance teams * Represent the team externally through benchmarks, whitepapers, and open-source contributions WHAT WE LOOK FOR * BS/MS/PhD in Computer Science, or a related field * Strong software engineering background (6+ years or equivalent) in performance-critical systems * Proven track record of owning complex system components and driving architectural decisions end-to-end * Deep understanding of ML inference internals: attention, MLPs, recurrent modules, quantization, sparse operations, etc. * Hands-on experience with CUDA, GPU programming, and key libraries (cuBLAS, cuDNN, NCCL, etc.) * Strong background in distributed systems design, including RPC frameworks, queuing, RPC batching, sharding, memory partitioning * Demonstrated ability to uncover and solve performance bottlenecks across layers (kernel, memory, networking, scheduler) * Experience building instrumentation, tracing, and profiling tools for ML models * Ability to lead through influence - work closely with ML researchers, translate novel model ideas into production systems * Excellent communication and leadership skills, with a proactive and ownership-driven mindset * Bonus: published research or open-source contributions in ML systems, inference optimization, or model serving Pay Range Transparency Databricks is committed to fair and equitable compensation practices. The pay range(s) for this role is listed below and represents the expected salary range for non-commissionable roles or on-target earnings for commissionable roles. Actual compensation packages are based on several factors that are unique to each candidate, including but not limited to job-related skills, depth of experience, relevant certifications and training, and specific work location. Based on the factors above, Databricks anticipates utilizing the full width of the range. The total compensation package for this position may also include eligibility for annual performance bonus, equity, and the benefits listed above. For more information regarding which range your location is in visit our page here. Local Pay Range $190,900—$232,800 USD About Databricks Databricks is the data and AI company. More than 10,000 organizations worldwide — including Comcast, Condé Nast, Grammarly, and over 50% of the Fortune 500 — rely on the Databricks Data Intelligence Platform to unify and democratize data, analytics and AI. Databricks is headquartered in San Francisco, with offices around the globe and was founded by the original creators of Lakehouse, Apache Spark™, Delta Lake and MLflow. To learn more, follow Databricks on Twitter, LinkedIn and Facebook. Benefits At Databricks, we strive to provide comprehensive benefits and perks that meet the needs of all of our employees. For specific details on the benefits offered in your region click here. Our Commitment to Diversity and Inclusion At Databricks, we are committed to fostering a diverse and inclusive culture where everyone can excel. We take great care to ensure that our hiring practices are inclusive and meet equal employment opportunity standards. Individuals looking for employment at Databricks are considered without regard to age, color, disability, ethnicity, family or marital status, gender identity or expression, language, national origin, physical and mental ability, political affiliation, race, religion, sexual orientation, socio-economic status, veteran status, and other protected characteristics. Compliance If access to export-controlled technology or source code is required for performance of job duties, it is within Employer's discretion whether to apply for a U.S. government license for such positions, and Employer may decline to proceed with an applicant on this basis alone.
P-1285 ABOUT THIS ROLE As a staff software engineer for GenAI Performance and Kernel, you will own the design, implementation, optimization, and correctness of the high-performance GPU kernels powering our GenAI inference stack. You will lead development of highly-tuned, low-level compute paths, manage trade-offs between hardware efficiency and generality, and mentor others in kernel-level performance engineering. You will work closely with ML researchers, systems engineers, and product teams to push the state-of-the-art in inference performance at scale. WHAT YOU WILL DO * Lead the design, implementation, benchmarking, and maintenance of core compute kernels (e.g. attention, MLP, softmax, layernorm, memory management) optimized for various hardware backends (GPU, accelerators) * Drive the performance roadmap for kernel-level improvements: vectorization, tensorization, tiling, fusion, mixed precision, sparsity, quantization, memory reuse, scheduling, auto-tuning, etc. * Integrate kernel optimizations with higher-level ML systems * Build and maintain profiling, instrumentation, and verification tooling to detect correctness, performance regressions, numerical issues, and hardware utilization gaps * Lead performance investigations and root-cause analysis on inference bottlenecks, e.g. memory bandwidth, cache contention, kernel launch overhead, tensor fragmentation * Establish coding patterns, abstractions, and frameworks to modularize kernels for reuse, cross-backend portability, and maintainability * Influence system architecture decisions to make kernel improvements more effective (e.g. memory layout, dataflow scheduling, kernel fusion boundaries) * Mentor and guide other engineers working on lower-level performance, provide code reviews, help set best practices * Collaborate with infrastructure, tooling, and ML teams to roll out kernel-level optimizations into production, and monitor their impact WHAT WE LOOK FOR * BS/MS/PhD in Computer Science, or a related field * Deep hands-on experience writing and tuning compute kernels (CUDA, Triton, OpenCL, LLVM IR, assembly or similar sort) for ML workloads * Strong knowledge of GPU/accelerator architecture: warp structure, memory hierarchy (global, shared, register, L1/L2 caches), tensor cores, scheduling, SM occupancy, etc. * Experience with advanced optimization techniques: tiling, blocking, software pipelining, vectorization, fusion, loop transformations, auto-tuning * Familiarity with ML-specific kernel libraries (cuBLAS, cuDNN, CUTLASS, oneDNN, etc.) or open kernels * Strong debugging and profiling skills (Nsight, NVProf, perf, vtune, custom instrumentation) * Experience reasoning about numerical stability, mixed precision, quantization, and error propagation * Experience in integrating optimized kernels into real-world ML inference systems; exposure to distributed inference pipelines, memory management, and runtime systems * Experience building high-performance products leveraging GPU acceleration * Excellent communication and leadership skills — able to drive design discussions, mentor colleagues, and make trade-offs visible * A track record of shipping performance-critical, high-quality production software * Bonus: published in systems/ML performance venues (e.g. MLSys, ASPLOS, ISCA, PPoPP), experience with custom accelerators or FPGA, experience with sparsity or model compression techniques Pay Range Transparency Databricks is committed to fair and equitable compensation practices. The pay range(s) for this role is listed below and represents the expected salary range for non-commissionable roles or on-target earnings for commissionable roles. Actual compensation packages are based on several factors that are unique to each candidate, including but not limited to job-related skills, depth of experience, relevant certifications and training, and specific work location. Based on the factors above, Databricks anticipates utilizing the full width of the range. The total compensation package for this position may also include eligibility for annual performance bonus, equity, and the benefits listed above. For more information regarding which range your location is in visit our page here. Local Pay Range $190,900—$232,800 USD About Databricks Databricks is the data and AI company. More than 10,000 organizations worldwide — including Comcast, Condé Nast, Grammarly, and over 50% of the Fortune 500 — rely on the Databricks Data Intelligence Platform to unify and democratize data, analytics and AI. Databricks is headquartered in San Francisco, with offices around the globe and was founded by the original creators of Lakehouse, Apache Spark™, Delta Lake and MLflow. To learn more, follow Databricks on Twitter, LinkedIn and Facebook. Benefits At Databricks, we strive to provide comprehensive benefits and perks that meet the needs of all of our employees. For specific details on the benefits offered in your region click here. Our Commitment to Diversity and Inclusion At Databricks, we are committed to fostering a diverse and inclusive culture where everyone can excel. We take great care to ensure that our hiring practices are inclusive and meet equal employment opportunity standards. Individuals looking for employment at Databricks are considered without regard to age, color, disability, ethnicity, family or marital status, gender identity or expression, language, national origin, physical and mental ability, political affiliation, race, religion, sexual orientation, socio-economic status, veteran status, and other protected characteristics. Compliance If access to export-controlled technology or source code is required for performance of job duties, it is within Employer's discretion whether to apply for a U.S. government license for such positions, and Employer may decline to proceed with an applicant on this basis alone.
Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. The Auth0 Team Auth0 is an easy-to-implement authentication and authorization platform designed by developers for developers. We make access to applications safe, secure, and seamless for the more than 100 million daily logins around the world. Our modern approach to identity enables this Tier-Ø global service to deliver convenience, privacy, and security so customers can focus on innovation. This team focuses on providing tenant-level protections to our customers, at scale. From bot detection to brute-force to suspicious IP throttling and beyond, this team often provides the first line of defense for Auth0 customers. The Staff Software Engineer At Okta, we’re building the next generation of authentication for the GenAI era. We’re looking for a Staff Software Engineer to join the AI DevEx team at Auth0. This role is pivotal in extending and complementing our Auth for GenAI offering by building the infrastructure, tooling, and developer experiences that empower both human developers and AI agents to build secure, intelligent applications. Auth0 Emerging Tech is the Engineering organization where we take care of the hottest technology out there: we ship fast, we don't break things. We are a dynamic and collaborative distributed and diverse team. We value ownership, learning and innovation. This is an ideal role for an engineer who enjoys building for other engineers, working across stacks, and shaping the future of AI enablement in production systems. You'll collaborate across engineering, product, and security teams to drive meaningful improvements in developer tooling, agent authentication, orchestration frameworks, and real-world demos. What you will be doing: * Design and Build Developer Tooling that helps developers secure and manage infrastructure like MCP servers * Build Demo Applications that showcase secure, identity-powered AI use cases in real-world environments * Contribute to Open Source Projects, both within Auth0 and across the broader AI + identity ecosystem * Write and Maintain High-Quality Documentation including API references, quickstarts, and best practices for both developers and AI-native tooling (e.g., llm.txt) * Drive Integration with Emerging AI Frameworks by creating adapters, utilities, and interfaces for agent runtimes and orchestration layers * Collaborate with Design, Product, and Security teams to align on developer needs, roadmap direction, and compliance requirements * Mentor and Support Other Engineers, setting strong examples in code quality, testing practices, and architectural thinking * Influence Engineering Standards by leading design discussions and contributing to team-wide architectural decisions * Ensure Resilience and Security of systems involved in agent-to-agent or model-to-service communication You Might Be a Good Fit If You * Experience in software engineering with a proven track record in building tools, frameworks, or platforms for other developers * Proficiency in JavaScript/TypeScript, Golang and/or Python, and the ability to move fluidly between front-end and back-end contexts * Experience working with LLM APIs, agent runtimes, orchestration layers, or prompt pipelines * Familiarity with authentication and authorization systems, especially standards like OAuth2, OIDC, and JWT * Demonstrated experience leading architecture and design efforts for scalable, production-grade systems * Comfort contributing to and maintaining open source projects and engaging with developer communities * A passion for documentation as part of the developer experience—not just writing code, but making it understandable and usable * Ability to thrive in highly collaborative environments with cross-functional stakeholders Technologies You May Work With * Languages: JavaScript, TypeScript, Python * Frameworks: React, Next.js, FastAPI * AI Ecosystem: Model APIs, orchestration runtimes, prompt management systems, agent toolkits * Auth0 Stack: Token Vault, Async Authorization, Fine-Grained Authorization (FGA) #LI-Hybrid P23578_3268067 Below is the annual base salary range for candidates located in San Francisco Bay Area. Your actual base salary will depend on factors such as your skills, qualifications, experience, and work location. In addition, Okta offers equity (where applicable), bonus, and benefits, including health, dental and vision insurance, 401(k), flexible spending account, and paid leave (including PTO and parental leave) in accordance with our applicable plans and policies. To learn more about our Total Rewards program please visit: https://rewards.okta.com/us. The annual base salary range for this position for candidates located in the San Francisco Bay area is between: $188,000—$282,000 USD The Okta Experience * Supporting Your Well-Being * Driving Social Impact * Developing Talent and Fostering Connection + Community We are intentional about connection. Our global community, spanning over 20 offices worldwide, is united by a drive to innovate. Your journey begins with an immersive, in-person onboarding experience designed to accelerate your impact and connect you to our mission and team from day one. Okta is an Equal Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, ancestry, marital status, age, physical or mental disability, or status as a protected veteran. We also consider for employment qualified applicants with arrest and convictions records, consistent with applicable laws. If reasonable accommodation is needed to complete any part of the job application, interview process, or onboarding please use this Form to request an accommodation. Notice for New York City Applicants & Employees: Okta may use Automated Employment Decision Tools (AEDT), as defined by New York City Local Law 144, that use artificial intelligence, machine learning, or other automated processes to assist in our recruitment and hiring process. In accordance with NYC Local Law 144, if you are an applicant or employee residing in New York City, please click here to view our full NYC AEDT Notice.