user avatar

Senior Software Engineer III – AI Infrastructure

Black Eagle Defense

Posted today

Job Requirements

Fort Meade, MD
Top Secret/SCI Full Scope Polygraph
Career Level not specified
$209,000 - $266,000

Job Description

Job Description

SALARY RANGE $209,000 - $266,000/year

DUTIES As a successful candidate for the Senior Software Engineer III - AI Infrastructure role, you will lead the development, operation, and evolution of the next generation of AI infrastructure that powers innovation across the customer organization. As a senior full-stack engineer, you will be responsible for designing, implementing, and maintaining critical AI platform capabilities that support scalable inference services and a growing ecosystem of AI-enabled applications. You will provide technical leadership across AI infrastructure initiatives, guiding the design and operation of reliable, secure, and high-performance platform components while collaborating with engineers, stakeholders, and integrated teams across the organization. In addition to hands-on engineering responsibilities, you will mentor team members, support professional development through coaching and feedback, and help coordinate technical efforts to deliver mission objectives. You will help establish technical standards, governance frameworks, and engineering best practices while driving adoption of emerging technologies that improve the scalability, reliability, and effectiveness of enterprise AI capabilities. Through leadership, systems engineering, cloud technologies, automation, and platform development, you will help shape the foundation that enables the customer's long-term AI strategy and operational success.

Required Skills

SKILLS
  • Design, implement, and optimize infrastructure supporting AI model inference at scale
  • Lead the development, deployment, and maintenance of production AI services and applications, including retrieval-augmented generation (RAG), autonomous agents, and emerging AI technologies
  • Serve as the technical lead for AI infrastructure initiatives, coordinating efforts across integrated engineering and platform teams
  • Provide leadership, mentorship, coaching, and professional development support to assigned team members
  • Conduct regular one-on-one meetings and provide constructive feedback to support team growth and performance
  • Act as the team point of contact for contract administration and operational coordination activities
  • Analyze complex and ambiguous requirements to define scalable, maintainable, and effective technical solutions
  • Establish and maintain technical standards, policies, governance frameworks, and engineering best practices
  • Drive the adoption of modern technologies, processes, and engineering methodologies across teams and organizations
  • Implement and oversee monitoring, logging, and observability capabilities to improve platform visibility and operational awareness
  • Ensure the availability, reliability, scalability, performance, and security of AI platform components and supporting infrastructure
  • Collaborate with engineers, stakeholders, and leadership to align technical solutions with organizational objectives
  • Communicate technical concepts, project status, risks, and recommendations to stakeholders at multiple organizational levels
  • Support the continuous improvement and modernization of AI infrastructure, cloud environments, and platform operations
  • Balance hands-on engineering responsibilities with leadership, coordination, and strategic planning activities
  • Contribute to the long-term growth and success of enterprise AI capabilities through technical leadership and platform innovation


QUALIFICATIONS Twelve (12) years' experience as a SWE in programs and contracts of similar scope, type, and complexity is required. A Bachelor's degree in Computer Science or a related discipline from an accredited college or university is required. Four (4) years of additional SWE experience on projects with similar software processes may be substituted for a bachelor's degree.

Additional Requirements:
  • Extensive experience designing, building, deploying, and operating large-scale production systems
  • Deep expertise integrating complex systems across diverse technologies, platforms, and environments
  • Hands-on experience with cloud engineering and solution deployment within AWS environments
  • Advanced proficiency in administering Kubernetes environments and implementing modern deployment patterns
  • Strong Python development skills for infrastructure, automation, and application development efforts
  • Experience implementing, scaling, and maintaining observability solutions using technologies such as APM, OpenTelemetry, Grafana, and Prometheus
  • Proven ability to lead technical initiatives and drive adoption of new technologies, processes, and engineering practices
  • Experience developing and implementing technical policies, standards, and governance frameworks
  • Strong understanding of cloud-native architectures, infrastructure automation, and operational excellence principles
  • Excellent communication, stakeholder management, and leadership skills across technical and non-technical audiences
  • Ability to balance hands-on engineering responsibilities with leadership, coordination, and strategic planning activities
  • Strong change management and organizational influence skills
  • Strong analytical, troubleshooting, and problem-solving abilities
  • Experience leading cross-functional teams and coordinating efforts across multiple stakeholders and organizations
  • Ability to operate effectively within ambiguous environments and establish structure for evolving requirements
  • Experience supporting the full lifecycle of cloud-native applications and platform services from design through production operations


Desired Skills

NICE-TO-HAVES
  • Experience with AI inference serving technologies such as vLLM, LiteLLM, or similar platforms
  • Experience developing solutions using agentic AI frameworks such as LangChain or comparable technologies
  • Knowledge of vector databases, embedding models, and semantic search architectures
  • Experience designing and supporting retrieval-augmented generation (RAG) solutions and AI-enabled applications
  • Familiarity with large language model deployment, optimization, and inference workflows
  • Experience with high-performance computing environments and distributed systems architectures
  • Knowledge of scalable data processing, distributed computing, and resource optimization techniques
  • Experience supporting enterprise AI platforms and machine learning infrastructure
  • Familiarity with emerging AI technologies, frameworks, and platform capabilities
  • Demonstrated track record of successfully driving technical transformation and organizational change initiatives
  • Experience influencing engineering culture, best practices, and technology adoption across teams
  • Experience leading modernization efforts involving cloud, platform, or AI infrastructure technologies
  • Ability to champion innovation while balancing operational stability and organizational objectives
group id: 91130336