Job Requirements
San Francisco, CA Omaha, NE Boston, MA Washington, DC
Secret Polygraph not specified
Mid Level Career (5+ yrs experience)
$170,000 - $225,000
Job Description
ClearanceJobs Worforce Solutions is seeking a Site Reliability Engineer – Infrastructure Operations for our client based in San Francisco, CA.
Our client is a cutting-edge startup delivering AI-based solutions for critical defense applications serving the U.S. Department of Defense and international defense customers. They operate 24/7 across multiple computing environments—cloud hyperscalers, on-premise GPU clusters, field-deployed systems, and supercomputing centers—with zero tolerance for downtime. The team is lean and mission-focused, with every member directly impacting the delivery of life-critical systems to their customers.
This is a 24/7 on-call infrastructure operations role focused on maintaining mission-critical AI systems at 99% uptime SLA. You'll be the primary responder for infrastructure incidents, monitoring system health, and optimizing existing production pipelines for reliability and performance.
This role is NOT about building new infrastructure from scratch. Our foundations are established. You'll be the operational backbone keeping them running reliably while continually optimizing for performance, cost, and customer SLA commitments.
Key Responsibilities
Primary Focus – On-Call Operations & Monitoring (60%)
● 24/7 on-call rotation as the primary responder for infrastructure incidents and production alerts
● Monitor system health and SLA metrics across all customer contracts; escalate to engineering team when needed
● Triage and respond to production incidents: analyze logs/metrics, diagnose root causes, execute remediation, and document post-mortems
● Build and maintain comprehensive observability: whitebox (application-level) and blackbox (system-level) monitoring
● Ensure Datadog and PagerDuty alerting strategies are tuned to catch issues without alert fatigue
● Develop and automate incident response playbooks to minimize MTTR (mean time to recovery)
● Maintain 99% uptime SLA across defense and government contracts
Secondary Focus – Infrastructure Optimization & Reliability (40%)
● Optimize existing infrastructure for performance, cost, and reliability (not redesign from scratch)
● Identify and address infrastructure bottlenecks through capacity planning and performance tuning
● Maintain ETL pipelines and data quality for core forecasting operations
● Collaborate with software engineers on deployment processes and CI/CD improvements
● Document runbooks, playbooks, and operational procedures for team scalability
What You'll Bring
Minimum Qualifications
● 5+ years of hands-on experience in SRE, infrastructure operations, or DevOps roles
● Python proficiency – it's our primary language; you'll write operational automation tools
● Deep experience with monitoring/observability platforms (Datadog, Prometheus, Grafana, Splunk, ELK, or equivalent)
● Proven on-call incident response experience
● Comfort with 24/7 on-call rotations – you understand production-critical operations and can escalate appropriately
● AWS or Google Cloud infrastructure experience (VPCs, load balancing, databases, serverless)
● Linux systems administration at a production level
● Active Secret clearance or ability to obtain one (required for this role)
Preferred Skills
● Infrastructure as Code (Terraform, CloudFormation) – nice to have but learnable on the job
● Kubernetes operations (EKS/GKE) – operational expertise more valuable than deep design
● Experience with message queues (SNS/SQS) and caching (Redis)
● GPU resource optimization and ML workload observability
● Experience supporting mission-critical systems for government/defense customers
● Background in incident response and blameless postmortem culture
What We're NOT Looking For
● Someone to redesign infrastructure from scratch (our foundations are solid)
● Pure CI/CD specialists focused on build systems (that's secondary here)
● Infrastructure architects without on-call operations experience
● Candidates uncomfortable with 24/7 on-call responsibilities
What You'll Get
● Competitive salary
● Small, world-class engineering team (5 engineers + CTO) – you'll have impact on every decision
● Mission-critical work for U.S. Air Force, Navy, and international defense customers
● Modern tech stack: Python, AWS/GCP, Datadog, PagerDuty, GitHub, Kubernetes
● Clear SRE culture: monitoring first, blameless postmortems, automation over heroics
● Bay Area location with relocation flexibility for exceptional candidates
The Reality of On-Call
This is a genuinely 24/7 role:
● You'll be on rotation with other engineers (currently 5-person team)
● Monitoring alerts notify you when systems degrade; you jump on them ASAP
● If you can't respond, it escalates to the next person in the rotation
● Customers are U.S. Air Force, Navy, and international defense agencies – downtime matters
● 99% SLA means ~7.2 hours of allowable downtime per month across all contracts
● This is not a role for someone who wants "off-the-grid" availability
The client prioritizes operational discipline over hero culture. If you're uncomfortable with this level of on-call responsibility, this role is not a fit.
Our client is a cutting-edge startup delivering AI-based solutions for critical defense applications serving the U.S. Department of Defense and international defense customers. They operate 24/7 across multiple computing environments—cloud hyperscalers, on-premise GPU clusters, field-deployed systems, and supercomputing centers—with zero tolerance for downtime. The team is lean and mission-focused, with every member directly impacting the delivery of life-critical systems to their customers.
This is a 24/7 on-call infrastructure operations role focused on maintaining mission-critical AI systems at 99% uptime SLA. You'll be the primary responder for infrastructure incidents, monitoring system health, and optimizing existing production pipelines for reliability and performance.
This role is NOT about building new infrastructure from scratch. Our foundations are established. You'll be the operational backbone keeping them running reliably while continually optimizing for performance, cost, and customer SLA commitments.
Key Responsibilities
Primary Focus – On-Call Operations & Monitoring (60%)
● 24/7 on-call rotation as the primary responder for infrastructure incidents and production alerts
● Monitor system health and SLA metrics across all customer contracts; escalate to engineering team when needed
● Triage and respond to production incidents: analyze logs/metrics, diagnose root causes, execute remediation, and document post-mortems
● Build and maintain comprehensive observability: whitebox (application-level) and blackbox (system-level) monitoring
● Ensure Datadog and PagerDuty alerting strategies are tuned to catch issues without alert fatigue
● Develop and automate incident response playbooks to minimize MTTR (mean time to recovery)
● Maintain 99% uptime SLA across defense and government contracts
Secondary Focus – Infrastructure Optimization & Reliability (40%)
● Optimize existing infrastructure for performance, cost, and reliability (not redesign from scratch)
● Identify and address infrastructure bottlenecks through capacity planning and performance tuning
● Maintain ETL pipelines and data quality for core forecasting operations
● Collaborate with software engineers on deployment processes and CI/CD improvements
● Document runbooks, playbooks, and operational procedures for team scalability
What You'll Bring
Minimum Qualifications
● 5+ years of hands-on experience in SRE, infrastructure operations, or DevOps roles
● Python proficiency – it's our primary language; you'll write operational automation tools
● Deep experience with monitoring/observability platforms (Datadog, Prometheus, Grafana, Splunk, ELK, or equivalent)
● Proven on-call incident response experience
● Comfort with 24/7 on-call rotations – you understand production-critical operations and can escalate appropriately
● AWS or Google Cloud infrastructure experience (VPCs, load balancing, databases, serverless)
● Linux systems administration at a production level
● Active Secret clearance or ability to obtain one (required for this role)
Preferred Skills
● Infrastructure as Code (Terraform, CloudFormation) – nice to have but learnable on the job
● Kubernetes operations (EKS/GKE) – operational expertise more valuable than deep design
● Experience with message queues (SNS/SQS) and caching (Redis)
● GPU resource optimization and ML workload observability
● Experience supporting mission-critical systems for government/defense customers
● Background in incident response and blameless postmortem culture
What We're NOT Looking For
● Someone to redesign infrastructure from scratch (our foundations are solid)
● Pure CI/CD specialists focused on build systems (that's secondary here)
● Infrastructure architects without on-call operations experience
● Candidates uncomfortable with 24/7 on-call responsibilities
What You'll Get
● Competitive salary
● Small, world-class engineering team (5 engineers + CTO) – you'll have impact on every decision
● Mission-critical work for U.S. Air Force, Navy, and international defense customers
● Modern tech stack: Python, AWS/GCP, Datadog, PagerDuty, GitHub, Kubernetes
● Clear SRE culture: monitoring first, blameless postmortems, automation over heroics
● Bay Area location with relocation flexibility for exceptional candidates
The Reality of On-Call
This is a genuinely 24/7 role:
● You'll be on rotation with other engineers (currently 5-person team)
● Monitoring alerts notify you when systems degrade; you jump on them ASAP
● If you can't respond, it escalates to the next person in the rotation
● Customers are U.S. Air Force, Navy, and international defense agencies – downtime matters
● 99% SLA means ~7.2 hours of allowable downtime per month across all contracts
● This is not a role for someone who wants "off-the-grid" availability
The client prioritizes operational discipline over hero culture. If you're uncomfortable with this level of on-call responsibility, this role is not a fit.
group id: ClearanceJobsSC