Job Requirements
Chantilly, VA
Public Trust Polygraph Unspecified
Career Level not specified
Salary not specified
Join Premium to unlock estimated salaries
Job Description
OVERVIEW:
We are seeking an ML Ops Engineer to own the machine-learning lifecycle in production. You will be responsible for getting the five detection models from trained artifact to live, low-latency serving, then keeping them healthy -monitored, versioned, and retrained. Your product is the models running well in production, not the data pipeline underneath them.
GENERAL DUTIES:
REQUIRED QUALIFICATIONS:
DESIRED QUALIFICATIONS:
CLEARANCE:
We are seeking an ML Ops Engineer to own the machine-learning lifecycle in production. You will be responsible for getting the five detection models from trained artifact to live, low-latency serving, then keeping them healthy -monitored, versioned, and retrained. Your product is the models running well in production, not the data pipeline underneath them.
GENERAL DUTIES:
- Model release management in MLflow - versioning, aliasing, promotion and rollback, champion/challenger across the five models.
- Serving models for real-time inference - package and optimize PyTorch models, run them in the low-latency inference workers, hold the <60s>
- Model and prediction monitoring with Evidently - data, concept, and prediction drift; performance decay; alerting - and closing the loop back to retraining.
- Automated retraining / continuous training - Airflow pipelines that retrain (including GPU training on EKS), validate against gates, and promote new model versions safely.
- Training/serving consistency - manage the Feast online/offline boundary to prevent training-serving skew.
- Reproducibility and governance - experiment tracking, model lineage/provenance, and model cards / approval gates for federal AI accountability.
REQUIRED QUALIFICATIONS:
- Owned the full production ML lifecycle - trained artifact to live serving to be monitored/retrained. Not model-building only, and not data-pipeline-building only.
- Model registry and experiment tracking - MLflow or equivalent (SageMaker, Weights & Biases, Vertex): versioning, promotion, rollback, lineage.
- Model serving for real-time/low-latency inference - embedded serving or a model server (TorchServe, Triton, KServe, Seldon, BentoML): model loading, optimization, latency debugging.
- Model and data drift monitoring - Evidently or equivalent; defining model-quality metrics and acting on decay.
- Automated retraining / CT pipelines and model CI/CD - validation gates, champion/challenger, shadow or canary rollouts for models.
- PyTorch (or TensorFlow) in production - packaging, optimizing (ONNX/quantization a plus), serving; debugging inference correctness and latency.
- Feature store consumption (Feast or equivalent) with real focus on training/serving skew.
- Kubernetes and Docker to package and deploy model workloads (Helm); Prometheus/Grafana for model and inference metrics.
- Strong Python and solid software engineering (tests, reproducibility) - not notebook-only.
DESIRED QUALIFICATIONS:
- The streaming pipeline you serve models into - Kafka + Bytewax (or Flink, Spark Streaming, Kafka Streams). You integrate with it; the data engineer owns it.
- Apache Airflow used specifically for ML orchestration (training, promotion, drift jobs).
- GPU training/serving on Kubernetes/EKS (CUDA/NVIDIA images).
- OpenShift and/or air-gapped model deployment.
- AWS GovCloud / FedRAMP / FIPS 140-2 / IL4-5, and federal AI governance - model cards, provenance, OSCAL, explainable scoring.
- Graph ML, autoencoders, and anomaly detection (our detection approach); security/behavioral feature work.
- Model artifacts in object storage (S3/MinIO); a warehouse (Redshift or equivalent) for offline evaluation data.
CLEARANCE:
- Active U.S. Citizenship with the eligibility to gain a clearance
group id: 90943786