Minimum qualifications:
- Bachelor's degree in Computer Science or similar technical field, or equivalent practical experience.
- 15 years of experience as a software engineer or 13 years with an advanced degree.
- Experience delivering large-scale capacity planning, IaaS/PaaS solutions, or fleet management systems.
Preferred qualifications:
- Master's or PhD in Computer Science or a field related to Networking or Security Systems.
- Experience leading and delivering large-scale, complex network infrastructure security initiatives from concept to production.
- Experience influencing and leading without direct authority, building relationships and driving alignment across unique infrastructure and product teams.
- Experience creating a compelling, shared vision that motivates engineering organizations, coupled with clear communication and a collaborative leadership approach.
- Exceptional communication skills, with the ability to articulate complex technical concepts to both engineering leaders and non-technical stakeholders.
About the job:
Google runs one of the largest computational fleets in the world, with compute, storage, networking and dedicated accelerators spread across all continents. As Principal Engineer, you will own the architectural outlook and end-to-end technical strategy for our global capacity planning ecosystem. You will serve as the primary technical authority for a challenge, knitting together complex planning constraints, optimization scenarios, and machine deployment workflows to get optimal capacity online in the shortest time.
As the Principal Engineer, Cloud Capacity Planning, you will bridge the gap between long-term product roadmaps and engineering execution. You will oversee the technical integrity of programs supporting everything from general purpose compute to specialized offerings, directly enabling massive-scale Cloud ML serving and training workloads. You will knit together the seams of different technical, product, and financial organizations to produce a globally optimal outcome for Cloud capacity amidst ongoing supply constraints and fluctuating enterprise demand.
In this role, you will utilize deep technical stewardship across forecasting, demand-supply matching, constraint-based planning, and scenario analysis to drive customer-focused, efficient, and sustainable resource ecosystem outcomes. You will deeply engage with and evolve our IaaS and PaaS tooling and processes to ensure they outpace the growth of Google Cloud's most impactful strategic programs.
Google Cloud accelerates every organization’s ability to digitally transform its business and industry. We deliver enterprise-grade solutions that leverage Google’s cutting-edge technology, and tools that help developers build more sustainably. Customers in more than 200 countries and territories turn to Google Cloud as their trusted partner to enable growth and solve their most critical business problems.
Individual pay is determined by factors including job-related skills, experience, and relevant education or training.US: $307000 - $427000 (USD) + 30% bonus target + equity + benefits
Learn more about benefits at Google.
Responsibilities:
- Identify, manage, and design mitigation strategies for technical and organizational risks that threaten delivery timelines.
- Define and drive the end-to-end architectural roadmap and technical strategy for the global capacity planning ecosystem, ensuring all major technical decisions align with Google-level OKRs and long-term business goals.
- Drive the development of automated IaaS and PaaS tooling, resource lifecycle management, resource forecasting, and real-time dashboards to improve end-to-end planning fidelity, track capacity health, and accelerate execution velocity.
- Act as the primary cross-functional technical bridge between capacity infrastructure, engineering organizations, and strategic product/financial partners to resolve complex dependencies and resource constraints.
- Lead the technical translation of enterprise customer business forecasts into specific infrastructure requirements across cores, storage, and specialized offerings (e.g., TPUs and advanced GPUs) to support both traditional computing and complex ML serving and training demands.
Learn more about this Employer on their Career Site
