Job Requirements
Tampa, FL
Top Secret/SCI Polygraph Unspecified
Career Level not specified
Salary not specified
Join Premium to unlock estimated salaries
Job Description
OVERVIEW:
We are seeking a Data Engineer with strong hands-on experience designing, developing, and managing large-scale data workflows across structured and unstructured datasets. This role focuses heavily on building reliable RAG (Retrieval-Augmented Generation) pipelines, orchestrating ETL/ELT processes, and deploying scalable data systems in AWS.
GENERAL DUTIES:
REQUIRED QUALIFICATIONS:
DESIRED QUALIFICATIONS:
CLEARANCE:
We are seeking a Data Engineer with strong hands-on experience designing, developing, and managing large-scale data workflows across structured and unstructured datasets. This role focuses heavily on building reliable RAG (Retrieval-Augmented Generation) pipelines, orchestrating ETL/ELT processes, and deploying scalable data systems in AWS.
GENERAL DUTIES:
- Design, build, and maintain RAG pipelines, including document ingestion, indexing, embedding workflows, and model retrieval flows.
- Develop and manage structured and unstructured data pipelines supporting analytics, ML, and application workloads. Build and optimize ETL/ELT pipelines in AWS using services such as S3, Lambda, Step Functions, EMR, Glue, ECS/EKS, and IAM best practices. Implement and operate NiFi flows for high-throughput, low-latency data ingestion and transformation.
- Develop, orchestrate, and schedule workflows using Prefect, ensuring reliability, observability, and proper error handling. Implement indexing, search, and retrieval patterns using ElasticSearch, including schema design, cluster management, and query optimization.
- Collaborate closely with architecture, ML, and application teams to support scalable data solutions.
- Ensure data quality, lineage, governance, and security across all pipelines.
- Monitor system performance and troubleshoot issues across distributed data systems.
REQUIRED QUALIFICATIONS:
- Solid understanding of ETL/ELT processes and data modeling best practices.
- Hands-on experience implementing workflows in Prefect (Prefect 2.0 preferred). In-depth knowledge of ElasticSearch indexing, cluster management, and search optimization.
- Proficiency in Python and familiarity with common data libraries (Pandas, PySpark, requests, etc.).
DESIRED QUALIFICATIONS:
- Strong experience with Apache NiFi for data flow management and real-time ingestion.
- Experience building or maintaining RAG pipelines (e.g., vector databases, embeddings, document chunking strategies, retrieval optimization).
- Proven ability to manage structured and unstructured data pipelines at scale.
- Experience with AWS cloud services for data engineering.
- Strong version control and CI/CD experience.
- Experience with vector databases (OpenSearch, Pinecone, Weaviate, etc.)
- Familiarity with containerized workflows (Docker, Kubernetes)
- Experience supporting LLM or generative AI production systems
- Background in distributed systems, streaming platforms, or data mesh architectures
CLEARANCE:
- Active TS/SCI clearance minimum required
group id: 90943786