Join Our Team as Senior DevOps Engineer (San Francisco)
Join a fast-growing, venture-backed startup building mission-critical grid technology that helps prevent power outages and wildfires. Backed by $100 million in investment from top firms like Sequoia Capital, Tiger Global, and Y Combinator, the company is scaling its product, customer base, and engineering organization with urgency.
As a Senior DevOps Engineer, you’ll own the cloud and Kubernetes foundation that keeps high-volume device telemetry, customer products, and engineering systems reliable. You’ll work across AWS, EKS, Kafka/MSK-style streaming, Terraform, GitOps, observability, security, cost, and incident response in a hands-on startup environment where senior engineers shape standards, reduce operational risk, and improve how teams ship. This is a high-ownership role responsible for the stability and scalability of critical IoT infrastructure for large energy companies.
Responsibilities
- Design, build, and maintain scalable, secure, and highly available infrastructure on AWS (EKS, EC2, RDS / Aurora Postgres, MSK, S3, VPC, IAM).
- Manage and optimize Kubernetes clusters (EKS) across multiple environments, and deploy applications using Argo CD with GitOps best practices.
- Implement and maintain CI/CD pipelines using GitHub Actions, including reusable workflows, build/push/scan flows for ECR, and frontend deployment pipelines.
- Operate and tune Kafka-based event streaming on Amazon MSK for high-throughput, low-latency device data pipelines.
- Define and manage Infrastructure as Code with Terraform and Terragrunt, with reusable modules, sensible environment separation, and review-friendly plans.
- Manage identity and access across platforms with Auth0 / EntraID integrations, IAM roles for service accounts (IRSA), and short-lived credentials.
- Build and maintain observability with Grafana, Loki, Prometheus / Mimir, and related tooling so on-call engineers can quickly find and fix issues.
- Monitor and optimize infrastructure cost across environments, partnering with engineering teams on right-sizing, capacity planning, and waste reduction.
- Partner with our Cloud Security team to enforce security standards, integrate with SIEM tooling, and respond to vulnerabilities and incidents.
- Debug complex production issues across infrastructure, deployment, and networking layers, and turn the lessons learned into automation and runbooks.
Requirements
- 5+ years in DevOps, SRE, or Platform Engineering with production experience operating AWS infrastructure.
- Deep hands-on experience administering Kubernetes (EKS or equivalent) and deploying via GitOps (Argo CD or Flux).
- Proficiency with Infrastructure as Code using Terraform; comfort with Terragrunt or a similar wrapper.
- Hands-on experience designing and maintaining CI/CD pipelines, preferably with GitHub Actions and reusable workflows.
- Production experience operating distributed systems such as Kafka (MSK).
- Strong understanding of networking, DNS, TLS, and security best practices, including IdP-driven access control (Auth0, EntraID, or similar).
- Solid experience with monitoring and logging stacks such as Grafana, Loki, Prometheus, Mimir, or equivalents.
- Ability to debug complex production issues across infrastructure, deployment, and networking layers.
- Comfortable working in Linux environments with strong scripting skills (Python or Bash preferred for automation).
- Knowledge of version control workflows, automated testing, and release management.
- You are a US Citizen or Permanent Resident
Bonus Skills
Your application will have a higher chance of standing out if you have one (or more) of the following skills or experiences. - Experience operating Apollo Router / GraphQL federation gateways in production. - Experience operating Argo Workflows or similar Kubernetes-native job / pipeline runners in production. - Familiarity with Databricks or ML Ops pipelines for data and model deployment. - Experience designing, operating, and exercising Disaster Recovery (DR) environments, including cross-region replication, backups, and tested failover runbooks. - Experience with Tailscale or other zero-trust networking tools. - Experience supporting IoT / embedded fleets at scale, including secure device-to-cloud connectivity. - Experience in high-growth startup environments where you must wear many hats.
Location & Schedule
- Onsite role based in San Francisco, CA.
- Relocation support provided.
Benefits
We offer competitive benefits that help employees to thrive, grow and enjoy their lives. These benefits include: - Health, Dental and Vision insurance, free parking and a commuter allowance - Stock option plan - Conveniently located office — directly across the street from Pleasant Hill BART, close to the highway with parking provided
Position Overview
- Location
- San Francisco
- Type
- Full-time
- Compensation
- $10,000 – $16,700/mo
Interested in this role?
Upload your resume and we'll be in touch within 2 business days.