Full-time Hiring Now

Senior Platform Engineer (San Francisco)

San Francisco, US $10,000 – $16,700/mo

Join Our Team as Senior Platform Engineer (San Francisco)

Join a fast-growing, venture-backed startup building mission-critical grid technology that helps prevent power outages and wildfires. Backed by $100 million in investment from top firms like Sequoia Capital, Tiger Global, and Y Combinator, the company is scaling its product, customer base, and engineering organization with urgency.

As a Senior Platform Engineer, you’ll treat internal developer experience as a product—discovering engineer pain points, defining paved paths, shaping platform strategy, and building self-service workflows that help teams move faster with confidence. You’ll work in a hands-on startup environment where priorities evolve, ownership is high, and your decisions directly influence how software teams build, deploy, and operate critical IoT infrastructure for large energy companies. This is a chance to create impact at both the product and platform level.

Responsibilities

  • Lead the design, build, and rollout of an internal developer platform on top of AWS, EKS, Argo CD, and GitHub Actions that lets engineers create, deploy, and operate services with minimal friction.
  • Own and evolve our service templates, Helm chart conventions, and Argo CD App-of-Apps patterns so that adding or migrating a service is a guided, low-risk experience.
  • Build and maintain reusable GitHub Actions workflows (build / push / scan, frontend build / deploy, SonarQube scans, semantic release) and improve CI feedback loops, build times, and caching.
  • Define and enforce platform standards for observability — structured logs into Loki, metrics into Prometheus / Mimir, dashboards in Grafana, and SLOs / alerts wired in by default.
  • Build self-service tooling around environments, secrets, feature flags, and access — so that the right thing is easy and the wrong thing is hard to do by accident.
  • Own the developer-facing aspects of identity and access (Auth0, IdP integrations, Tailscale access, IRSA / service accounts) and keep onboarding and offboarding smooth.
  • Partner with DevOps on infrastructure changes, with Cloud Security on guardrails, and with backend / frontend / data / firmware teams to understand their pain points and prioritize platform investments.
  • Mentor engineers across the org on platform conventions, lead design reviews for new services, and push back on patterns that don’t scale.
  • Treat the platform as a product: gather feedback, define roadmaps, write docs, and measure adoption and reliability.

Requirements

  • 5+ years in Platform Engineering, DevOps, or SRE roles, including significant experience building and shipping developer-facing tooling for other engineering teams.
  • Track record of owning and delivering platform initiatives end-to-end, from design through adoption, with limited day-to-day supervision.
  • Strong working knowledge of Kubernetes (EKS or similar) and GitOps workflows with Argo CD or Flux.
  • Hands-on experience with Infrastructure as Code using Terraform; comfort with Terragrunt or a similar wrapper.
  • Solid experience with CI/CD systems, ideally GitHub Actions, including reusable / composable workflows and release automation.
  • Working knowledge of AWS core services (EKS, EC2, RDS, S3, IAM, VPC, ECR) and how to compose them into reliable, secure platforms.
  • Experience designing developer abstractions — Helm charts, service templates, internal CLIs, scaffolding tools, or Backstage-style portals — that other engineering teams easily interact with.
  • Strong programming skills in Python, Bash, or TypeScript for building tooling and automation.
  • Experience integrating observability (Grafana, Loki, Prometheus / Mimir, OpenTelemetry, or similar) as a default rather than an afterthought.
  • Strong written communication skills, with a habit of writing docs, runbooks, and wikis that engineers can actually use.

Bonus Skills

Your application will have a higher chance of standing out if you have one (or more) of the following skills or experiences. - Experience building or operating Apollo Router / GraphQL federation gateways and supporting subgraph development workflows. - Experience with Backstage or a comparable internal developer portal. - Experience integrating Argo Workflows or similar Kubernetes-native job / pipeline runners into a developer platform. - Familiarity with Databricks or ML Ops pipelines and the developer experience around data / model deployment. - Experience with Tailscale, Auth0, EntraID, or other identity / zero-trust networking tooling. - Familiarity with cloud architectures supporting IoT / embedded systems and distributed, low-power devices. - Experience in high-growth startup environments where you must wear many hats. - You are a US Citizen or Permanent Resident.

Location & Schedule

  • Onsite role based in San Francisco, CA.
  • Relocation support provided.

Benefits

Join a fast-growing, venture-backed startup building mission-critical technology for a safer, more resilient U.S. power grid. Backed by more than $100 million in investment, the company is scaling its product, infrastructure, and customer deployments with urgency. As a Senior DevOps Engineer, you’ll own the cloud and Kubernetes foundation that keeps high-volume device telemetry, customer products, and engineering systems reliable. You’ll work across AWS, EKS, Kafka/MSK-style streaming, Terraform, GitOps, observability, security, cost, and incident response in a hands-on startup environment where senior engineers shape standards, reduce operational risk, and improve how teams ship. This is a high-ownership role for someone who wants production impact tied to real-world infrastructure.

About This Role

Own production cloud operations across AWS, EKS, Kubernetes, Argo CD, Terraform/Terragrunt, GitHub Actions, Kafka/MSK-style streaming, Aurora Postgres, and Grafana/Loki/Prometheus observability. You’ll improve reliability, networking, incident response, security guardrails, cost efficiency, disaster recovery, and CI/CD workflows for high-volume telemetry systems powering critical grid monitoring products.

Responsibilities

  • Design, build, and maintain scalable, secure, and highly available infrastructure on AWS (EKS, EC2, RDS / Aurora Postgres, MSK, S3, VPC, IAM).
  • Manage and optimize Kubernetes clusters (EKS) across multiple environments, and deploy applications using Argo CD with GitOps best practices.
  • Implement and maintain CI/CD pipelines using GitHub Actions, including reusable workflows, build/push/scan flows for ECR, and frontend deployment pipelines.
  • Operate and tune Kafka-based event streaming on Amazon MSK for high-throughput, low-latency device data pipelines.
  • Define and manage Infrastructure as Code with Terraform and Terragrunt, with reusable modules, sensible environment separation, and review-friendly plans.
  • Manage identity and access across platforms with Auth0 / EntraID integrations, IAM roles for service accounts (IRSA), and short-lived credentials.
  • Build and maintain observability with Grafana, Loki, Prometheus / Mimir, and related tooling so on-call engineers can quickly find and fix issues.
  • Monitor and optimize infrastructure cost across environments, partnering with engineering teams on right-sizing, capacity planning, and waste reduction.
  • Partner with our Cloud Security team to enforce security standards, integrate with SIEM tooling, and respond to vulnerabilities and incidents.
  • Debug complex production issues across infrastructure, deployment, and networking layers, and turn the lessons learned into automation and runbooks.

Requirements

  • 5+ years in DevOps, SRE, or Platform Engineering with production experience operating AWS infrastructure.
  • Deep hands-on experience administering Kubernetes (EKS or equivalent) and deploying via GitOps (Argo CD or Flux).
  • Proficiency with Infrastructure as Code using Terraform; comfort with Terragrunt or a similar wrapper.
  • Hands-on experience designing and maintaining CI/CD pipelines, preferably with GitHub Actions and reusable workflows.
  • Production experience operating distributed systems such as Kafka (MSK).
  • Strong understanding of networking, DNS, TLS, and security best practices, including IdP-driven access control (Auth0, EntraID, or similar).
  • Solid experience with monitoring and logging stacks such as Grafana, Loki, Prometheus, Mimir, or equivalents.
  • Ability to debug complex production issues across infrastructure, deployment, and networking layers.
  • Comfortable working in Linux environments with strong scripting skills (Python or Bash preferred for automation).
  • Knowledge of version control workflows, automated testing, and release management.

Bonus Skills

Your application will have a higher chance of standing out if you have one (or more) of the following skills or experiences. - Experience operating Apollo Router / GraphQL federation gateways in production. - Experience operating Argo Workflows or similar Kubernetes-native job / pipeline runners in production. - Familiarity with Databricks or ML Ops pipelines for data and model deployment. - Experience designing, operating, and exercising Disaster Recovery (DR) environments, including cross-region replication, backups, and tested failover runbooks. - Experience with Tailscale or other zero-trust networking tools. - Experience supporting IoT / embedded fleets at scale, including secure device-to-cloud connectivity. - Experience in high-growth startup environments where you must wear many hats.

Location & Schedule

  • Onsite role based in San Francisco, CA.
  • Relocation support provided.

Benefits

We offer competitive benefits that help employees to thrive, grow and enjoy their lives. These benefits include: - Health, Dental and Vision insurance, free parking and a commuter allowance - Stock option plan - Conveniently located office — directly across the street from Pleasant Hill BART, close to the highway with parking provided

Key Skills

  • WS
  • EKS
  • Argo CD

Position Overview

Location
San Francisco, US
Type
Full-time
Compensation
$10,000 – $16,700/mo

Interested in this role?

Upload your resume and we'll be in touch within 2 business days.

Apply — Senior Platform Engineer (San Francisco)

Upload your resume (PDF or Word). We'll review and reach out.