Job Requirements
Chantilly, VA
Top Secret Polygraph not specified
Mid Level Career (5+ yrs experience)
$230,000 - $250,000
Job Description
LLMOps Engineer | Top Secret Clearance required - Chantilly VA
Salary to 250K and great benefits!
Are you passionate about building and operating production-scale Large Language Model infrastructure? We're looking for an experienced LLMOps Engineer to help deploy, optimize, and scale self-hosted AI platforms.
What we're looking for:
✅ Deep understanding of LLM internals (Transformers, tokenization, inference, quantization, fine-tuning)
✅ Hands-on experience with vLLM deployment and production operations
✅ Kubernetes, Docker, GPU clusters, and AWS, Azure, or GCP
✅ Expertise serving open-weight models like Llama, Mistral, and Qwen
✅ Experience with tensor parallelism, continuous batching, PagedAttention, KV-cache optimization, and model quantization (AWQ, GPTQ, FP8)
✅ Strong DevOps background with CI/CD, Infrastructure as Code, observability, monitoring, and incident response
✅ Proven success managing production LLM environments for performance, reliability, scalability, and cost efficiency
If you've been responsible for keeping a live, self-hosted LLM platform running at peak performance, we'd love to hear from you.
📩 Contact Debbie at The Josef Group, dpeda@thejosefgroup.com to learn more about this exciting opportunity.
Salary to 250K and great benefits!
Are you passionate about building and operating production-scale Large Language Model infrastructure? We're looking for an experienced LLMOps Engineer to help deploy, optimize, and scale self-hosted AI platforms.
What we're looking for:
✅ Deep understanding of LLM internals (Transformers, tokenization, inference, quantization, fine-tuning)
✅ Hands-on experience with vLLM deployment and production operations
✅ Kubernetes, Docker, GPU clusters, and AWS, Azure, or GCP
✅ Expertise serving open-weight models like Llama, Mistral, and Qwen
✅ Experience with tensor parallelism, continuous batching, PagedAttention, KV-cache optimization, and model quantization (AWQ, GPTQ, FP8)
✅ Strong DevOps background with CI/CD, Infrastructure as Code, observability, monitoring, and incident response
✅ Proven success managing production LLM environments for performance, reliability, scalability, and cost efficiency
If you've been responsible for keeping a live, self-hosted LLM platform running at peak performance, we'd love to hear from you.
📩 Contact Debbie at The Josef Group, dpeda@thejosefgroup.com to learn more about this exciting opportunity.
group id: 10111992