Job Requirements
MD
Top Secret/SCI Full Scope Polygraph
Career Level not specified
$260,000 - $285,000
Job Description
Applicants must already hold a TS/SCI clearance with a current Full Scope Polygraph before applying. We are not sponsoring or processing clearances for this position. Candidates who are seeking to obtain a clearance do not meet the minimum requirements.
Salary Range: $260k-$285k per year with an additional $65k-$71k in immediately vested company 401(k) contributions
Description: You will act as a Software Engineer with a focus on DevOps, platform engineering, and infrastructure automation for a high-impact mission delivering AI/ML capabilities to end users. Your work will center on building reliable, scalable infrastructure and deployment pipelines rather than application development, ensuring continuous availability of mission-critical services through infrastructure-as-code, containerization, and orchestration. You will implement comprehensive observability and monitoring solutions that track system health and user usage patterns. Collaborating with data scientists and system administrators, you will own the platform layer that enables rapid, resilient delivery of AI technologies to the mission.
Responsibilities:
Skills Requirements:
Nice to Haves:
YOE Requirement: 12 yrs., B.S. in a technical discipline or 4 additional yrs. in place of B.S.
Salary Range: $260k-$285k per year with an additional $65k-$71k in immediately vested company 401(k) contributions
Description: You will act as a Software Engineer with a focus on DevOps, platform engineering, and infrastructure automation for a high-impact mission delivering AI/ML capabilities to end users. Your work will center on building reliable, scalable infrastructure and deployment pipelines rather than application development, ensuring continuous availability of mission-critical services through infrastructure-as-code, containerization, and orchestration. You will implement comprehensive observability and monitoring solutions that track system health and user usage patterns. Collaborating with data scientists and system administrators, you will own the platform layer that enables rapid, resilient delivery of AI technologies to the mission.
Responsibilities:
- Deploy and manage containerized AI/ML services using Kubernetes in on-premises environments.
- Implement observability and monitoring solutions using Prometheus, Grafana, and related tooling to track system health and usage metrics.
- Build and maintain CI/CD pipelines and infrastructure-as-code to enable reliable, automated deployment of platform services.
- Design and implement resilient infrastructure patterns to ensure high availability and performance of mission-critical capabilities.
- Collaborate with stakeholders to translate operational and reliability requirements in platform engineering solutions.
Skills Requirements:
- Experience deploying and managing services in on-premises infrastructure environments.
- Service containerization and orchestration with Docker/Kubernetes in production environments.
- Experience with observability and monitoring tools such as Prometheus, Grafana, or equivalent platforms.
- Knowledge of CI/CD pipelines, GitOps workflows.
- Service containerization and deployment with Docker/Kubernetes.
- Experience with version control (git).
- Atlassian Tools (Jira, Confluence).
Nice to Haves:
- Experience with model serving technologies (vLLM, TorchServe, TensorFlow Serving, Triton Inference Server).
- Knowledge of AI infrastructure tools and frameworks (LiteLLM, Bifrost, Ray, Ollama).
- Knowledge of service mesh technologies (Istio, Linkerd) or advanced Kubernetes patterns.
- Awareness of current AI/ML landscape including emerging models, capabilities, and applicability to different use cases.
- Knowledge of end-to-end SIGINT collection and analysis systems.
- Experience with production CNO capabilities and operations.
YOE Requirement: 12 yrs., B.S. in a technical discipline or 4 additional yrs. in place of B.S.
group id: 91132822