Staff Software Engineer - Cloud Platform Products (Agentic Workloads)
Budget / Salary$120,000–190,000
TypeFull-time job
LocationUnited States
Posted4 hours ago
With a career at The Home Depot, you can be yourself and also be part of something bigger.
Position Purpose
As a Staff Software Engineer, you will be the technical leader accountable for the architecture, design, and delivery of cloud platform products that enable and support agentic workloads across The Home Depot. These products serve engineering teams building, deploying, and operating AI-powered applications, and non-engineering teams who need secure, governed self-service experiences to build and run their own agentic solutions.
You will turn platform strategy into scalable technical solutions and then make sure they ship. You own technical delivery outcomes, not just designs. You will act as a force multiplier: you raise the delivery speed, quality, and independence of the engineers and teams around you through hands-on leadership, clear technical direction, reusable patterns, and removing blockers. You will lead by influence, partnering closely with your Software Engineer Manager, Platform Reliability Engineering, Security, Architecture, and consuming teams to speed up an AI-driven product development lifecycle.
Key Responsibilities
40% Delivery Leadership & Force Multiplication
Owns technical delivery of cloud platform products supporting agentic workloads, from design through release and ongoing operation, in partnership with the Software Engineer Manager
Breaks ambiguous initiatives into well-scoped, sequenced work that engineers can execute on independently, and keeps delivery predictable, incremental, and low-risk
Identifies and removes technical blockers, dependencies, and bottlenecks across the team and partner teams before they affect delivery
Raises team throughput by building reusable services, libraries, templates, paved paths, and automation that cut repeated effort and toil
Sets up and continuously improves the team’s delivery practices, including CI/CD, progressive delivery, IaC/GitOps (Terraform, Helm, Argo CD), testing strategy, and release standards
Champions and models AI-assisted development and agentic workflows to measurably speed up how the team designs, builds, tests, and ships
Contributes hands-on code on the most complex, highest-risk, or most critical-path work, and pairs with engineers to unblock and level them up
Tracks delivery health (lead time, deployment frequency, change failure rate, recovery time) and drives improvements
Drives adoption through clear documentation, reference architectures, self-service workflows, and developer experience improvements for both technical and non-technical users
30% Architecture, Technical Leadership & Strategy
Defines and drives the technical vision, architecture, and long-term evolution of cloud platform products supporting agentic workloads
Leads the design of secure, scalable, resilient platform capabilities on Google Cloud and GKE that serve both engineering and non-engineering users
Establishes architectural standards, engineering patterns, and governance guardrails (security, compliance, cost) that make safe self-service the default
Leads design reviews, architecture reviews, and technical deep dives, and participates in enterprise review boards to drive consistency
Partners with Enterprise Architecture, Security, and Platform Reliability Engineering to ensure solutions meet enterprise standards
Leads evaluation and adoption of emerging technologies in AI infrastructure, LLM gateways, agent runtimes, MCP/tool integration, model serving, and orchestration
Serves as a trusted technical advisor to the Software Engineer Manager and product leadership on roadmap feasibility, sequencing, and technical risk
15% Reliability, Support & Operational Excellence
Partners with Platform Reliability Engineering to define SLOs, observability, alerting, runbooks, and on-call readiness for every product before and after launch
Serves as the technical escalation point and leads technical response for complex production incidents
Drives root cause analysis and makes sure systemic fixes are prioritized and delivered
Designs platforms with supportability, scalability, and cost efficiency built in from the start
Supports highly available 24x7 retail operations
15% Mentorship & Engineering Excellence
Mentors and develops engineers through technical coaching, design guidance, code reviews, and pairing, with the explicit goal of growing their ability to deliver independently
Raises engineering standards for quality, maintainability, security, and operational excellence across the team
Delegates meaningful technical ownership to grow other engineers into technical leaders
Fosters communities of practice in cloud-native development, agentic AI, and platform engineering
Advocates for continuous learning and hands-on experimentation with emerging technologies
Direct Manager / Direct Reports
Reports to Software Engineer Manager
0 direct reports
Travel Requirements
Typically requires overnight travel less than 10% of the time
Minimum Qualifications
Must be eighteen years of age or older
Must be legally permitted to work in the United States
Mastery of an object-oriented programming language (preferably Go)
Extensive experience designing, building, and delivering distributed cloud-native systems
Proficiency with Google Cloud Platform and Kubernetes (GKE) in a production environment
Experience leading architecture, technical decisions, and delivery across multiple applications or teams
Preferred Qualifications
8+ years of software engineering experience with demonstrated technical leadership across large-scale platforms or distributed systems
Proven track record of delivering complex platform initiatives end to end and measurably improving the delivery speed and quality of surrounding engineers and teams
Deep expertise in Golang for cloud-native services, APIs, platforms, and developer tooling
Expert-level knowledge of Google Cloud (GKE, IAM, Networking, Cloud Run, Pub/Sub, Cloud Storage, Load Balancing, Security)
Expert-level experience with Kubernetes and container orchestration at scale
Deep expertise in IaC, GitOps, and automation (Terraform, Helm, Argo CD)
Extensive experience designing and running modern CI/CD and progressive delivery pipelines
Experience building platforms for agentic AI workloads: LLM gateways, agent runtimes, model serving, MCP integration, orchestration layers, vector databases
Experience designing developer platforms, APIs, and self-service experiences for both technical and non-technical users
Strong observability experience: OpenTelemetry, Prometheus, Grafana, tracing, logging, SLOs, alerting
Experience partnering with SRE / Reliability Engineering to build highly available production systems
Strong understanding of cloud security, identity, secrets management, network policy, and governance controls
Experience leading enterprise architecture reviews and influencing technical direction across teams
Demonstrated ability to align engineering, product, architecture, security, and operations stakeholders
Proven ability to mentor engineers and raise technical excellence across an organization
Experience supporting highly available 24x7 retail platforms
AI-driven mindset, using AI-assisted development and agentic workflows to improve engineering productivity and platform capabilities
Education
Minimum: the knowledge, skills, and abilities typically acquired through a bachelor’s degree program or equivalent in a related field. Preferred: no additional education.
Minimum Years of Work Experience: 7
Preferred Years of Work Experience: 8+
Minimum / Preferred Leadership Experience: None (technical leadership is covered under qualifications)
Certifications: None (Google Cloud Professional certifications are a plus)
Competencies
Drives Results
Plans and Aligns
Develops Talent
Collaborates
Manages Ambiguity
Cultivates Innovation
Communicates Effectively
Situational Adaptability
Nimble Learning
Interpersonal Savvy
For California, Colorado, Connecticut, Rhode Island, Nevada, New York City, Ithaca (NY), Westchester County (NY), and Washington residents:
The pay range for this position is between $120,000.00 - $190,000.00Originally posted on Himalayas
Position Purpose
As a Staff Software Engineer, you will be the technical leader accountable for the architecture, design, and delivery of cloud platform products that enable and support agentic workloads across The Home Depot. These products serve engineering teams building, deploying, and operating AI-powered applications, and non-engineering teams who need secure, governed self-service experiences to build and run their own agentic solutions.
You will turn platform strategy into scalable technical solutions and then make sure they ship. You own technical delivery outcomes, not just designs. You will act as a force multiplier: you raise the delivery speed, quality, and independence of the engineers and teams around you through hands-on leadership, clear technical direction, reusable patterns, and removing blockers. You will lead by influence, partnering closely with your Software Engineer Manager, Platform Reliability Engineering, Security, Architecture, and consuming teams to speed up an AI-driven product development lifecycle.
Key Responsibilities
40% Delivery Leadership & Force Multiplication
Owns technical delivery of cloud platform products supporting agentic workloads, from design through release and ongoing operation, in partnership with the Software Engineer Manager
Breaks ambiguous initiatives into well-scoped, sequenced work that engineers can execute on independently, and keeps delivery predictable, incremental, and low-risk
Identifies and removes technical blockers, dependencies, and bottlenecks across the team and partner teams before they affect delivery
Raises team throughput by building reusable services, libraries, templates, paved paths, and automation that cut repeated effort and toil
Sets up and continuously improves the team’s delivery practices, including CI/CD, progressive delivery, IaC/GitOps (Terraform, Helm, Argo CD), testing strategy, and release standards
Champions and models AI-assisted development and agentic workflows to measurably speed up how the team designs, builds, tests, and ships
Contributes hands-on code on the most complex, highest-risk, or most critical-path work, and pairs with engineers to unblock and level them up
Tracks delivery health (lead time, deployment frequency, change failure rate, recovery time) and drives improvements
Drives adoption through clear documentation, reference architectures, self-service workflows, and developer experience improvements for both technical and non-technical users
30% Architecture, Technical Leadership & Strategy
Defines and drives the technical vision, architecture, and long-term evolution of cloud platform products supporting agentic workloads
Leads the design of secure, scalable, resilient platform capabilities on Google Cloud and GKE that serve both engineering and non-engineering users
Establishes architectural standards, engineering patterns, and governance guardrails (security, compliance, cost) that make safe self-service the default
Leads design reviews, architecture reviews, and technical deep dives, and participates in enterprise review boards to drive consistency
Partners with Enterprise Architecture, Security, and Platform Reliability Engineering to ensure solutions meet enterprise standards
Leads evaluation and adoption of emerging technologies in AI infrastructure, LLM gateways, agent runtimes, MCP/tool integration, model serving, and orchestration
Serves as a trusted technical advisor to the Software Engineer Manager and product leadership on roadmap feasibility, sequencing, and technical risk
15% Reliability, Support & Operational Excellence
Partners with Platform Reliability Engineering to define SLOs, observability, alerting, runbooks, and on-call readiness for every product before and after launch
Serves as the technical escalation point and leads technical response for complex production incidents
Drives root cause analysis and makes sure systemic fixes are prioritized and delivered
Designs platforms with supportability, scalability, and cost efficiency built in from the start
Supports highly available 24x7 retail operations
15% Mentorship & Engineering Excellence
Mentors and develops engineers through technical coaching, design guidance, code reviews, and pairing, with the explicit goal of growing their ability to deliver independently
Raises engineering standards for quality, maintainability, security, and operational excellence across the team
Delegates meaningful technical ownership to grow other engineers into technical leaders
Fosters communities of practice in cloud-native development, agentic AI, and platform engineering
Advocates for continuous learning and hands-on experimentation with emerging technologies
Direct Manager / Direct Reports
Reports to Software Engineer Manager
0 direct reports
Travel Requirements
Typically requires overnight travel less than 10% of the time
Minimum Qualifications
Must be eighteen years of age or older
Must be legally permitted to work in the United States
Mastery of an object-oriented programming language (preferably Go)
Extensive experience designing, building, and delivering distributed cloud-native systems
Proficiency with Google Cloud Platform and Kubernetes (GKE) in a production environment
Experience leading architecture, technical decisions, and delivery across multiple applications or teams
Preferred Qualifications
8+ years of software engineering experience with demonstrated technical leadership across large-scale platforms or distributed systems
Proven track record of delivering complex platform initiatives end to end and measurably improving the delivery speed and quality of surrounding engineers and teams
Deep expertise in Golang for cloud-native services, APIs, platforms, and developer tooling
Expert-level knowledge of Google Cloud (GKE, IAM, Networking, Cloud Run, Pub/Sub, Cloud Storage, Load Balancing, Security)
Expert-level experience with Kubernetes and container orchestration at scale
Deep expertise in IaC, GitOps, and automation (Terraform, Helm, Argo CD)
Extensive experience designing and running modern CI/CD and progressive delivery pipelines
Experience building platforms for agentic AI workloads: LLM gateways, agent runtimes, model serving, MCP integration, orchestration layers, vector databases
Experience designing developer platforms, APIs, and self-service experiences for both technical and non-technical users
Strong observability experience: OpenTelemetry, Prometheus, Grafana, tracing, logging, SLOs, alerting
Experience partnering with SRE / Reliability Engineering to build highly available production systems
Strong understanding of cloud security, identity, secrets management, network policy, and governance controls
Experience leading enterprise architecture reviews and influencing technical direction across teams
Demonstrated ability to align engineering, product, architecture, security, and operations stakeholders
Proven ability to mentor engineers and raise technical excellence across an organization
Experience supporting highly available 24x7 retail platforms
AI-driven mindset, using AI-assisted development and agentic workflows to improve engineering productivity and platform capabilities
Education
Minimum: the knowledge, skills, and abilities typically acquired through a bachelor’s degree program or equivalent in a related field. Preferred: no additional education.
Minimum Years of Work Experience: 7
Preferred Years of Work Experience: 8+
Minimum / Preferred Leadership Experience: None (technical leadership is covered under qualifications)
Certifications: None (Google Cloud Professional certifications are a plus)
Competencies
Drives Results
Plans and Aligns
Develops Talent
Collaborates
Manages Ambiguity
Cultivates Innovation
Communicates Effectively
Situational Adaptability
Nimble Learning
Interpersonal Savvy
For California, Colorado, Connecticut, Rhode Island, Nevada, New York City, Ithaca (NY), Westchester County (NY), and Washington residents:
The pay range for this position is between $120,000.00 - $190,000.00Originally posted on Himalayas
Apply on Himalayas →
Job sourced from Himalayas. Applications happen directly on the original platform — we never collect your data.