Job Requirements
Herndon, VA
Top Secret/SCI Full Scope Polygraph
Career Level not specified
Salary not specified
Join Premium to unlock estimated salaries
Job Description
Opportunity via QSSHire | Talent-as-a-Service Recruiting
Role: Lead Site Reliability Engineer (SRE)
Location: Herndon, VA (Hybrid)
Clearance: TS/SCI with Full Scope Polygraph
Rate: Open
Are You a Lead Site Reliability Engineer Ready to Drive Reliability, Automation, and
Scalable Infrastructure?
QSSHire is seeking an experienced The Lead Site Reliability Engineer (SRE) will help
maintain and improve the reliability, performance, and security of software products and
platforms. This role combines operational excellence with engineering rigor, with a strong
emphasis on automation and Infrastructure-as-Code.
The Lead SRE will be a core part of the team responsible for ensuring uptime and system
health across services operating on both internal and client networks. The position will also
contribute to technical development, platform enhancements, compliance initiatives, and
the development of automated solutions for preventing, detecting, and remediating
operational issues.
In This Role, You’ll:
• Maintain the uptime, performance, and security of software products across
internal and client environments.
• Leverage automation and Infrastructure-as-Code to manage and scale
infrastructure.
• Respond to and resolve support requests during assigned SRE shifts.
• Escalate support requests as necessary to meet established Service Level
Agreements (SLAs).
• Develop SRE Playbooks designed to automate the avoidance, detection, and
remediation of operational issues.
• Participate in code reviews with a focus on quality and security.
• Contribute to accreditation and compliance initiatives.
• Support platform capability enhancements and other technical development
activities.
• Help maintain system health and operational reliability across production services.
• Support 24/7 production services through scheduled SRE duty rotations and shared
after-hours on-call coverage.
You’ll Succeed Here If You Have:
• Demonstrated foundation in automation and Infrastructure-as-Code.
• Demonstrated foundation in DevOps practices.
• Demonstrated understanding of distributed systems.
• Demonstrated experience working within cloud environments.
• Ability to maintain and improve the reliability, performance, security, uptime, and
overall health of software products and platforms.
• Ability to manage and scale infrastructure through automation and Infrastructure
as-Code.
• Ability to respond to and resolve operational support requests and escalate issues
as necessary to meet SLAs.
• Experience working with technologies and environments that include: Kubernetes,
Terraform, AWS, Go and/or Python
• Active U.S. Government security clearance.
• Ability to comply with applicable security obligations and procedural requirements.
Desired Skills:
• Experience developing SRE playbooks for the automated avoidance, detection, or
remediation of operational issues.
• Experience performing code reviews for quality and security purposes.
• Experience supporting accreditation and compliance initiatives.
• Experience contributing to platform capability enhancements and other technical
development activities.
• Experience supporting 24/7 production services.
• Experience supporting software products and platforms across both internal and
client environments.
Why QSSHire?
As a modern Talent-as-a-Service recruiting partner, QSSHire connects exceptional
professionals with impactful opportunities supporting mission-critical programs. We focus
on aligning specialized expertise with meaningful work while providing transparency,
flexibility, and long-term career growth.
Role: Lead Site Reliability Engineer (SRE)
Location: Herndon, VA (Hybrid)
Clearance: TS/SCI with Full Scope Polygraph
Rate: Open
Are You a Lead Site Reliability Engineer Ready to Drive Reliability, Automation, and
Scalable Infrastructure?
QSSHire is seeking an experienced The Lead Site Reliability Engineer (SRE) will help
maintain and improve the reliability, performance, and security of software products and
platforms. This role combines operational excellence with engineering rigor, with a strong
emphasis on automation and Infrastructure-as-Code.
The Lead SRE will be a core part of the team responsible for ensuring uptime and system
health across services operating on both internal and client networks. The position will also
contribute to technical development, platform enhancements, compliance initiatives, and
the development of automated solutions for preventing, detecting, and remediating
operational issues.
In This Role, You’ll:
• Maintain the uptime, performance, and security of software products across
internal and client environments.
• Leverage automation and Infrastructure-as-Code to manage and scale
infrastructure.
• Respond to and resolve support requests during assigned SRE shifts.
• Escalate support requests as necessary to meet established Service Level
Agreements (SLAs).
• Develop SRE Playbooks designed to automate the avoidance, detection, and
remediation of operational issues.
• Participate in code reviews with a focus on quality and security.
• Contribute to accreditation and compliance initiatives.
• Support platform capability enhancements and other technical development
activities.
• Help maintain system health and operational reliability across production services.
• Support 24/7 production services through scheduled SRE duty rotations and shared
after-hours on-call coverage.
You’ll Succeed Here If You Have:
• Demonstrated foundation in automation and Infrastructure-as-Code.
• Demonstrated foundation in DevOps practices.
• Demonstrated understanding of distributed systems.
• Demonstrated experience working within cloud environments.
• Ability to maintain and improve the reliability, performance, security, uptime, and
overall health of software products and platforms.
• Ability to manage and scale infrastructure through automation and Infrastructure
as-Code.
• Ability to respond to and resolve operational support requests and escalate issues
as necessary to meet SLAs.
• Experience working with technologies and environments that include: Kubernetes,
Terraform, AWS, Go and/or Python
• Active U.S. Government security clearance.
• Ability to comply with applicable security obligations and procedural requirements.
Desired Skills:
• Experience developing SRE playbooks for the automated avoidance, detection, or
remediation of operational issues.
• Experience performing code reviews for quality and security purposes.
• Experience supporting accreditation and compliance initiatives.
• Experience contributing to platform capability enhancements and other technical
development activities.
• Experience supporting 24/7 production services.
• Experience supporting software products and platforms across both internal and
client environments.
Why QSSHire?
As a modern Talent-as-a-Service recruiting partner, QSSHire connects exceptional
professionals with impactful opportunities supporting mission-critical programs. We focus
on aligning specialized expertise with meaningful work while providing transparency,
flexibility, and long-term career growth.
group id: 91142086