Post Job Free
Sign in

Senior Data Engineer - AWS & Knowledge Graphs

Location:
Longmont, CO
Salary:
90000
Posted:
September 30, 2026

Contact this candidate

Resume:

RAGHAVENDRA JAKKULA

Fort Collins, CO 970-***-**** ***************@*****.*** raghavendrajakkula.com linkedin.com/in/raghavendrajakkula3

SUMMARY

Senior Data Engineer with 5+ years of experience building enterprise-scale data platforms on AWS — designing Python and PySpark pipelines using AWS Glue, EMR, Kinesis, Firehose, Redshift, and S3 to support provider, member, claims, and network analytics. Deep expertise in graph database technologies including AWS Neptune (Gremlin traversal, SPARQL, RDF and property graph models) and knowledge graph construction pipelines applicable to Neo4j and TigerGraph patterns — building entity resolution, relationship modeling, and network traversal systems for healthcare analytics. Proven track record building AI-driven data solutions including LLM-powered knowledge graph enrichment, RAG pipelines, and intelligent data automation. Strong background in PII/PHI-aware data handling, HIPAA-adjacent security controls, audit trails, and enterprise data governance. Python and SQL expert. AWS Certified Data Engineer and Solutions Architect. CERTIFICATIONS

AWS Data Engineer Associate AWS Solutions Architect Associate AWS Cloud Practitioner Azure DevOps Engineer Expert AZ-400 (all valid 2025-2028)

TECHNICAL SKILLS

AWS Services: Redshift, S3, Glue, EMR, Kinesis, Firehose, Lambda, CloudWatch, IAM, DynamoDB, Neptune, EKS, Step Functions, MWAA — enterprise data lake and warehouse architectures, healthcare analytics data platforms, multi-account security Graph Databases: AWS Neptune (Gremlin traversal, SPARQL, RDF and property graph models), knowledge graph design and construction, entity resolution, relationship modeling; graph patterns applicable to Neo4j (Cypher) and TigerGraph — provider network graphs, member-claims graphs, healthcare network analytics Python and PySpark: Python (expert), PySpark — AWS Glue PySpark jobs, crawlers, Data Catalog, Glue Studio; Spark workload profiling, DAG analysis, executor tuning, partition skew optimization; enterprise-scale ETL/ELT pipeline development ETL and Pipelines: End-to-end ETL/ELT pipeline design, batch and real-time streaming, Apache Iceberg (ACID, time-travel, schema evolution), Airflow MWAA, Kinesis Firehose — provider, member, claims, and network data ingestion and processing Knowledge Graphs and AI: LangChain, LlamaIndex, RAG pipelines, LLM-powered knowledge graph enrichment, entity extraction and linking, graph-aware retrieval, MLflow — AI-driven data solutions and intelligent automation for healthcare analytics Healthcare Data: Provider, member, claims, and network data processing; PII/PHI-aware data handling, HIPAA-adjacent security controls, audit trails, data lineage tracking, access controls — compliance-ready healthcare data products Non-Relational and Graph DBs: DynamoDB (key-value), MongoDB (document), S3 object storage, AWS Neptune (graph), column-family databases — multi-store integration for diverse healthcare data types and relationship-centric analytics DevOps and CI/CD: Jenkins, Terraform, GitHub Actions, Docker, Kubernetes, GitOps — CI/CD pipelines, infrastructure as code, zero-downtime rollouts, automated quality gates, Linux administration (Ubuntu, RHEL) EXPERIENCE

Senior Data Engineer — Healthcare Analytics and Knowledge Graphs Aivanta Tech Inc., Remote Jan 2025 - Present

• Designed and built enterprise-scale Python and PySpark data pipelines on AWS using Glue, EMR, Kinesis, Redshift, S3, and Lambda — integrated structured and unstructured data from diverse upstream sources into graph and warehouse layers; built AWS Neptune property graph models for entity traversal and network analytics across provider, member, and operational relationship data.

• Constructed knowledge graph enrichment pipelines using LangChain and LlamaIndex — designed LLM-powered entity extraction, relationship linking, and graph population workflows; implemented Gremlin traversal queries and SPARQL endpoints for graph-aware retrieval; built RAG systems enabling natural language queries over graph-structured healthcare-adjacent data.

• Built AI-driven data automation frameworks using LangChain and MLflow eliminating manual data acquisition, validation, and publishing workflows — automated knowledge graph construction pipelines; deployed MLflow model serving and GenAI inference endpoints on EKS; reduced MTTD by 35% and MTTR by 30% through intelligent automation.

• Enforced PII/PHI-aware data handling and HIPAA-adjacent security standards — implemented IAM least-privilege, VPC security policies, audit logging, and data lineage tracking across all managed pipelines; instrumented with CloudWatch metrics, SLO dashboards, and data quality gates; partnered with ML and security teams preventing multiple major production outages.

• Executed DevOps/CI/CD automation (Jenkins, Terraform, GitHub Actions) with zero-downtime Kubernetes deployments; designed modular Terraform IaC framework reducing environment provisioning substantially; mentored engineers on AWS Glue PySpark best practices and Neptune graph database design patterns. Data Engineer and Software Engineer SAN R&D; Business Solutions, Fort Collins, CO Jan 2024 - Jan 2025

• Built enterprise ETL/ELT pipelines using Python, PySpark, and AWS Glue integrating multiple upstream systems into Redshift, Snowflake, and S3 — designed graph-compatible data models enabling relationship traversal analytics; implemented automated data quality gates and schema validation reducing data errors substantially.

• Automated CI/CD pipelines (Jenkins, Terraform, GitHub Actions) with zero-downtime Kubernetes rolling deployments — delivered self-healing monitoring with CloudWatch alarms and Lambda auto-remediation reducing MTTR substantially; designed modular Terraform IaC framework cutting environment provisioning from days to hours.

• Applied product analytics and measurement frameworks to operational data — defined key metrics, monitored trends, identified anomalies, and produced structured data summaries informing engineering and business decision-making across the platform team.

Software Engineer — Distributed Platform and Data Engineering Colorado State University, Fort Collins, CO Jan 2024 - Jan 2025

• Architected AWS distributed data platform (S3, Glue, EMR, Redshift, Spark) ingesting 50+ heterogeneous structured and unstructured sources — built PySpark transformation jobs and AWS Glue crawlers maintaining Data Catalog; achieved consistently high pipeline reliability with automated quality monitoring and anomaly detection on Linux (RHEL).

• Delivered multi-tenant platform features through full lifecycle — design, code review, containerized CI/CD, zero-downtime Kubernetes rollouts; tuned Spark executor memory, parallelism, and dynamic allocation on EMR eliminating resource contention; enforced IAM least-privilege, access-control audits, and secrets-management standards; translated operational data into insights for planning teams across dozens of stakeholders.

Software Engineer — Big Data and Healthcare Data Pipelines MedPlus Health Services, Hyderabad, India May 2020 - Dec 2023

• Built high-throughput batch and streaming ingestion pipelines in Python and PySpark on AWS (S3, Glue, EMR, Redshift, Kinesis, Lambda) processing large volumes of daily healthcare operational records — performed Spark workload profiling including DAG analysis, stage-level metrics, and I/O patterns; deployed on Kubernetes via Jenkins CI/CD on Linux improving throughput substantially and reducing costs.

• Implemented PII/PHI-aware data handling and security controls across all managed pipelines — deployed IAM role boundaries, VPC security group policies, audit logging, and data lineage tracking maintaining zero critical security incidents; designed ETL/ELT processes integrating relational and non-relational stores (DynamoDB, MongoDB, S3) across a large retail pharmacy network.

• Diagnosed recurring failures in multi-tenant Spark pipelines and implemented tiered automated alerting reducing escalations substantially — built graph-compatible data models for network relationship analytics; optimized DynamoDB and MongoDB access patterns under high-concurrency load reducing database latency significantly. PROJECTS

EvOne — AI-Powered Knowledge Graph Application Solo Developer evone-six.vercel.app 2024 - Present

• Built EvOne as a solo developer — a live consumer mobile app using LLM-powered intelligence to extract entities, relationships, and actionable knowledge from video content; designed knowledge graph construction pipelines using LangChain and graph-compatible data models for entity linking and relationship traversal.

• Built graph-aware RAG retrieval system enabling natural language queries over structured entity-relationship data — designed all data architecture, built ETL pipelines, and implemented data quality checks independently; shipped to production at evone-six.vercel.app demonstrating end-to-end graph data engineering.

• Applied data engineering principles throughout — designed graph entity extraction pipeline, built retrieval and traversal system, integrated LLMs to power knowledge graph enrichment, and shipped a working product available for real users today. Enterprise Knowledge Graph and AWS Glue Platform — Healthcare Analytics Python · PySpark · AWS Glue · AWS Neptune · Redshift · S3 · Apache Iceberg · LangChain · Airflow · Terraform

• Architected enterprise data platform with knowledge graph layer on AWS Neptune — built Python and PySpark AWS Glue pipelines ingesting provider, member, and network data; implemented Gremlin traversal queries, SPARQL endpoints, and graph construction workflows linking entities and relationships across analytics domains.

• Built LLM-powered knowledge graph enrichment using LangChain — automated entity extraction, relationship linking, and graph population from structured and unstructured data; implemented Apache Iceberg for ACID transactions and schema evolution; built full observability stack reducing incident response time substantially. EDUCATION

M.S., Computer Information Systems — Colorado State University, Fort Collins, CO (2025) Relevant Coursework: Applied Machine Learning, Statistical Methods for Data Analysis, Predictive Modeling, Regression Analysis, Big Data Systems

B.S., Computer Science — Osmania University, Hyderabad, India



Contact this candidate