Job Requirements
Remote
Top Secret Polygraph not specified
Career Level not specified
Salary not specified
Join Premium to unlock estimated salaries
Job Description
Company: Peraton (DCSA)
Job Title: AI Engineer
Location: Remote
Clearance: Active Top Secret
Top Skills
[8]+ years building applied ML/AI or data systems, with demonstrated delivery of LLM and RAG systems you personally built - not notebook demos.
Hands-on Databricks.
Document processing at scale: OCR, layout-aware parsing, chunking tradeoffs, poor-quality source handling.
Local/self-hosted LLM serving - vLLM, TGI, Ollama, llama.cpp, or equivalent - including running open-weight models in an isolated or air-gapped environment without reliance on external API endpoints.
Structured extraction and grounded generation with source attribution.
LLM evaluation methodology - you can explain how you measured correctness and what the evaluation missed.
Privacy-preserving synthetic data generation from CUI, PII, or comparably restricted source data, with an understanding of re-identification risk.
Strong Python.
Government or defense contracting experience.
This role builds LLM and document-intelligence capabilities on a greenfield data and AI platform in a high-trust federal environment. Work is hands-on, scoped to a near-term product demonstration, and built to migrate - documented, portable, and transferred to the internal team. Engagement is 1 year with option to extend. Active Top Secret clearance required.
Responsibilities
Build RAG and document-processing pipelines on the Databricks lakehouse - ingestion, OCR of mixed-quality sources, chunking, embedding, and retrieval.
Build LLM workflows for summarization, structured extraction, and evidence-grounded generation with source attribution.
Generate synthetic document corpora with the fidelity and quality variation needed for meaningful results.
Stand up the evaluation harness - retrieval quality, groundedness and hallucination checks, structured-output validity, human-in-the-loop review. Report results in numbers.
Package deliverables as jobs and Asset Bundles, tracked in MLflow, and document everything the internal team needs to own the work.
Required Qualifications
U.S. citizenship and active T5/SSBI federally adjudicated clearance required.
[8]+ years building applied ML/AI or data systems, with demonstrated delivery of LLM and RAG systems you personally built - not notebook demos.
Hands-on Databricks.
Document processing at scale: OCR, layout-aware parsing, chunking tradeoffs, poor-quality source handling.
Local/self-hosted LLM serving - vLLM, TGI, Ollama, llama.cpp, or equivalent - including running open-weight models in an isolated or air-gapped environment without reliance on external API endpoints.
Structured extraction and grounded generation with source attribution.
LLM evaluation methodology - you can explain how you measured correctness and what the evaluation missed.
Privacy-preserving synthetic data generation from CUI, PII, or comparably restricted source data, with an understanding of re-identification risk.
Strong Python.
Government or defense contracting experience.
Preferred Qualifications
RAG built inside a government or FedRAMP-authorized environment (Azure OpenAI in GCC High, AWS GovCloud, Bedrock within an authorized boundary).
Direct experience with FedRAMP Moderate, NIST 800-171, CMMC L2, or CUI handling.
Databricks Vector Search, Mosaic AI Agent Framework and Agent Evaluation, Asset Bundles, MLflow.
Unity Catalog governance.
H2O (h2oGPTe, Driverless AI).
Soft Skills
Self-directed execution against a fixed milestone with minimal oversight.
Honest reporting of model behavior - comfortable saying what an evaluation does and does not establish.
Collaboration across technical and non-technical teams.
Clear documentation and active knowledge transfer.
Thanks and Regards
Murali Sharma
202.828.3494
Murali@Nastechglobal.com
Job Title: AI Engineer
Location: Remote
Clearance: Active Top Secret
Top Skills
[8]+ years building applied ML/AI or data systems, with demonstrated delivery of LLM and RAG systems you personally built - not notebook demos.
Hands-on Databricks.
Document processing at scale: OCR, layout-aware parsing, chunking tradeoffs, poor-quality source handling.
Local/self-hosted LLM serving - vLLM, TGI, Ollama, llama.cpp, or equivalent - including running open-weight models in an isolated or air-gapped environment without reliance on external API endpoints.
Structured extraction and grounded generation with source attribution.
LLM evaluation methodology - you can explain how you measured correctness and what the evaluation missed.
Privacy-preserving synthetic data generation from CUI, PII, or comparably restricted source data, with an understanding of re-identification risk.
Strong Python.
Government or defense contracting experience.
This role builds LLM and document-intelligence capabilities on a greenfield data and AI platform in a high-trust federal environment. Work is hands-on, scoped to a near-term product demonstration, and built to migrate - documented, portable, and transferred to the internal team. Engagement is 1 year with option to extend. Active Top Secret clearance required.
Responsibilities
Build RAG and document-processing pipelines on the Databricks lakehouse - ingestion, OCR of mixed-quality sources, chunking, embedding, and retrieval.
Build LLM workflows for summarization, structured extraction, and evidence-grounded generation with source attribution.
Generate synthetic document corpora with the fidelity and quality variation needed for meaningful results.
Stand up the evaluation harness - retrieval quality, groundedness and hallucination checks, structured-output validity, human-in-the-loop review. Report results in numbers.
Package deliverables as jobs and Asset Bundles, tracked in MLflow, and document everything the internal team needs to own the work.
Required Qualifications
U.S. citizenship and active T5/SSBI federally adjudicated clearance required.
[8]+ years building applied ML/AI or data systems, with demonstrated delivery of LLM and RAG systems you personally built - not notebook demos.
Hands-on Databricks.
Document processing at scale: OCR, layout-aware parsing, chunking tradeoffs, poor-quality source handling.
Local/self-hosted LLM serving - vLLM, TGI, Ollama, llama.cpp, or equivalent - including running open-weight models in an isolated or air-gapped environment without reliance on external API endpoints.
Structured extraction and grounded generation with source attribution.
LLM evaluation methodology - you can explain how you measured correctness and what the evaluation missed.
Privacy-preserving synthetic data generation from CUI, PII, or comparably restricted source data, with an understanding of re-identification risk.
Strong Python.
Government or defense contracting experience.
Preferred Qualifications
RAG built inside a government or FedRAMP-authorized environment (Azure OpenAI in GCC High, AWS GovCloud, Bedrock within an authorized boundary).
Direct experience with FedRAMP Moderate, NIST 800-171, CMMC L2, or CUI handling.
Databricks Vector Search, Mosaic AI Agent Framework and Agent Evaluation, Asset Bundles, MLflow.
Unity Catalog governance.
H2O (h2oGPTe, Driverless AI).
Soft Skills
Self-directed execution against a fixed milestone with minimal oversight.
Honest reporting of model behavior - comfortable saying what an evaluation does and does not establish.
Collaboration across technical and non-technical teams.
Clear documentation and active knowledge transfer.
Thanks and Regards
Murali Sharma
202.828.3494
Murali@Nastechglobal.com
group id: 91142412