Post Job Free
Sign in

Data engineer

Location:
Allen, TX
Salary:
150000
Posted:
August 20, 2026

Contact this candidate

Resume:

ARATI TARPARA

LEAD DATA ENGINEER AZURE DATABRICKS PYSPARK SNOWFLAKE DATA ARCHITECTURE

Allen, TX 75013 469-***-**** **************@*****.*** linkedin.com/in/arati-tarpara US Citizen

PROFESSIONAL SUMMARY

Lead Data Engineer with 15+ years of combined experience in enterprise data engineering, ETL/ELT, cloud platforms, database development, healthcare business systems, and business intelligence across insurance, healthcare, and manufacturing. Experienced leading large-scale insurance carrier data platforms using Snowflake, Python, Apache Airflow, Azure, ADF, ADLS Gen2, Nexla, and UiPath. Strong hands-on background in advanced SQL, stored procedures, reusable pipeline frameworks, data quality, reconciliation, production troubleshooting, CI/CD, and month-end processing. Led 400+ production pipelines supporting 200+ insurance carriers and helped reduce carrier processing timelines from 45 days to 14 days. Experience maps directly to modern Lakehouse and Medallion patterns, including raw/landing, standardized, cleansed, and business-ready data layers. Familiar with Azure Databricks, PySpark, Delta Lake, Synapse Analytics, and healthcare interoperability concepts including EHR data, FHIR, HL7, HIPAA-aware data handling, medical billing, coding, and insurance verification. Recent hands-on POC exposure includes Snowflake Cortex and Streamlit for Snowflake-native AI and natural-language analytics use cases.

TECHNICAL SKILLS

Cloud & Data Platforms: Snowflake, Microsoft Azure (Azure Data Factory, ADLS Gen2, Azure SQL, Synapse Analytics), Azure Databricks (working knowledge/exposure), AWS S3

Lakehouse & Big Data: PySpark/Spark concepts, Delta Lake, Lakehouse Architecture, Medallion Architecture, structured and unstructured data processing, scalable batch processing

Data Engineering & Orchestration: Apache Airflow (Astronomer), ETL/ELT, Nexla, Informatica PowerCenter 10.2, SSIS, UiPath RPA, Streamlit, Files.com

Programming & SQL: Python (Pandas, NumPy), SQL, T-SQL, PL/SQL, DDL/DML, stored procedures, functions, triggers, views, Bash, R, VBA

Snowflake & Data Quality: Snowflake SQL, stored procedures, layered transformations, audit logging, validation, reconciliation, duplicate controls, root-cause analysis, SLA monitoring

DevOps: Azure DevOps, Git, branch management, PR-triggered releases, DAG deployment, CI/CD, Dev-to-Prod promotion

Healthcare Data: EHR/EMR concepts, claims and billing data, FHIR/HL7 concepts, insurance verification, HIPAA-aware data handling, patient and clinical workflow data, medical claims lifecycle, professional/institutional claims concepts, ICD-10, CPT, HCPCS familiarity, eligibility, reimbursement, denials, payer/provider data, billed/allowed/paid amounts, deductible, copay, coinsurance

AI & Analytics Enablement: Snowflake Cortex (POC/exposure), Streamlit, LLM-enabled analytics concepts, Claude, Microsoft Copilot, Power BI, Tableau, predictive analytics concepts

Domain: Insurance carrier and commission data, policy/insured attributes, premium calculations, agent hierarchy, payment frequency, healthcare billing and insurance verification

PROFESSIONAL EXPERIENCE

Integrity Marketing Group LLC Dallas-Fort Worth, TX (Hybrid) Oct 2021 - Present

Senior Data Engineer Lead Data Engineer

• Lead end-to-end ingestion and processing of 10,000+ structured and unstructured files monthly from 200+ insurance carrier sources into Snowflake using Airflow, Python, Nexla, UiPath RPA, ADF, and custom processing modules.

• Directed architecture changes that reduced carrier data processing timelines from 45 days to 14 days, improving month-end reconciliation and SLA performance.

• Designed and supported reusable Snowflake transformation patterns using stored procedures, standardized tables and views, validation rules, duplicate controls, audit logging, and carrier-specific business logic.

• Own production troubleshooting and root-cause analysis across ingestion and transformation pipelines using logs, quarantine/error data, job-level tracking, reconciliation results, and Snowflake query analysis.

• Partner with business analysts, finance teams, and carrier stakeholders to translate commission feed requirements and source-to-target mappings into scalable technical solutions.

• Manage Azure DevOps CI/CD processes for DAGs and data pipelines, including branch management, code review, controlled releases, and Dev-to-Prod deployment.

• Mentor and guide distributed engineering teams, promoting reusable development standards, code quality, documentation, and production support practices.

• Led month-end reconciliation improvements and achieved full reconciliation within two weeks, including carrier mapping updates, baseline changes, exception resolution, and production validation.

• Built hands-on Snowflake Cortex and Streamlit POC work to explore natural-language interaction with governed Snowflake data and AI-enabled analytics; this work is exploratory rather than production ML model development.

• Designed pipelines around layered/Medallion data architecture principles, separating landing/raw ingestion, standardized processing, cleansed data, and business-ready outputs to improve traceability, reprocessing, and downstream reliability.

• Process both structured and unstructured source files and apply reusable Python-based parsing, standardization, schema validation, business-rule validation, metadata handling, and duplicate detection before downstream warehouse processing.

• Use Azure Data Factory and ADLS Gen2 for cloud ingestion and storage patterns and work with Azure Synapse concepts for analytics and warehousing use cases; experience is directly transferable to Databricks/Delta Lake Lakehouse implementations.

• Participate in technical design reviews and architecture decisions covering orchestration, storage, transformation, performance, data quality, failure handling, observability, scalability, and production support.

• Automate operational controls such as file-count reconciliation, missing-file detection, quarantine checks, failure logging, and Jira-based exception workflows to improve reliability and reduce manual monitoring.

• Apply data-quality controls throughout the pipeline including record/file count validation, required-column checks, duplicate controls, source-to-target reconciliation, exception quarantine, and audit logging before data reaches downstream consumers.

Lead Data Engineer ETL Pipeline Architect Oct 2021 - Oct 2024

• Architected a scalable, carrier-agnostic ETL and automation framework using Snowflake, Nexla, UiPath, Azure, and Airflow patterns, allowing new carrier feeds and formats to be onboarded through reusable components.

• Oversaw delivery of 400+ production pipelines with 1,000+ sub-pipeline branches supporting analytics, reconciliation, and operational reporting across 200+ insurance carriers in 44 states.

• Established ADF and ADLS Gen2 integrations for carrier file ingestion and reusable mapping documentation, strengthening data validation and reducing manual development effort.

• Led replacement of a third-party reconciliation service with an in-house UiPath + Snowflake solution, reducing external dependency and improving auditability and governance.

HOBI International Inc Dallas, TX Nov 2019 - Oct 2021

Data Analyst Contract

• Built Informatica PowerCenter 10.2 ETL pipelines to ingest Avro files from AWS S3, reconcile data with PostgreSQL, and deliver cleansed datasets to Snowflake for reporting.

• Designed Power BI dashboards for Sales, Production, and executive teams using DirectQuery, Live Connection, DAX measures, and calculated columns.

• Created and optimized PL/SQL stored procedures, functions, and packages and performed database performance tuning using indexing and query optimization techniques.

• Created GIS layers in ArcGIS Pro and ArcGIS Online and published spatial analytics to Power BI.

Mar-Lan Industries Irving, TX Jun 2013 - Nov 2019

Data Analyst Full-Time

• Designed SQL Server database objects including tables, stored procedures, views, triggers, and user-defined functions supporting OLAP and enterprise reporting.

• Developed and optimized SSIS/DTS ETL packages across heterogeneous source systems and authored complex T-SQL and MDX queries.

• Tuned SQL performance using execution plans, clustered/non-clustered indexes, and database monitoring; supported backup/restore, reconciliation, and database administration across 10+ environments.

Shivom USA Inc. Denton, TX Nov 2011 - Jun 2013

Business Consultant Self-Employed

• Supported hospital medical billing and coding processes and assisted with accurate patient billing workflows.

• Performed insurance verification activities to support reimbursement processing and reduce avoidable claim issues.

• Assisted with patient documentation, pharmacy technician duties, medication inventory, and healthcare administrative operations.

• Worked with healthcare operational data and workflows involving patient, insurance, billing, coding, and clinical documentation, providing foundational experience for modern EHR/claims-data engineering environments.

• Applied healthcare privacy and data-handling awareness when working with patient and insurance information; familiar with HIPAA expectations and healthcare data confidentiality requirements.

• Supported healthcare billing and insurance workflows involving patient, provider, payer, diagnosis, procedure, eligibility, and reimbursement information, giving direct exposure to the business context behind medical claims data.

• Familiar with the medical claims lifecycle from eligibility and claim submission through acceptance/rejection, adjudication, payment, and denial follow-up, including common claim attributes such as service date, diagnosis/procedure codes, billed amount, allowed amount, paid amount, deductible, copay, and coinsurance.

• Working familiarity with ICD-10, CPT, and HCPCS coding concepts and how coded clinical and billing information is used in claims processing, reimbursement, reporting, and healthcare analytics.

• Understand professional and institutional claims concepts, payer/provider relationships, duplicate/invalid claim checks, denial analysis, and HIPAA/PHI-aware handling of patient and insurance data.

SELECTED PROJECTS

Enterprise Carrier Data Platform - Snowflake + Airflow

• Directed a reusable dual-DAG orchestration pattern with dynamic task expansion and parallel Snowflake stored-procedure execution across STAGE and CLEANSED layers.

• Implemented transformation logic covering policy and insured attributes, status data, premium calculations, agent hierarchy, and payment frequency, with per-carrier/job audit logging.

• Designed automated file movement from carrier portals through Files.com and Streamlit into ADLS Gen2, followed by validation, duplicate checking, downstream Snowflake processing, and exception handling.

Azure Lakehouse / Databricks Alignment

• Current architecture uses Azure cloud storage, layered data processing, Python transformations, orchestration, data-quality controls, and CI/CD patterns that align closely with Databricks Lakehouse and Bronze/Silver/Gold design principles.

• Working knowledge of PySpark/Spark DataFrame processing, Delta Lake concepts, partitioning, scalable transformations, schema enforcement/evolution, and performance-oriented Lakehouse design; positioned as transferable experience rather than overstated production Databricks ownership.

Healthcare Data & Interoperability Exposure

• Healthcare domain background includes medical billing, coding, insurance verification, patient documentation, and healthcare administrative workflows.

• Familiar with modern healthcare integration concepts including EHR/EMR data, FHIR, HL7, claims/billing datasets, HIPAA-aware handling, and analytics use cases that combine clinical and operational data.

Snowflake Cortex / Streamlit AI POC

• Explored a Snowflake-native AI use case for natural-language access to governed carrier processing and reconciliation data.

• Focused on business context, reliable underlying data, user interaction through Streamlit, and how Cortex capabilities can support business users without requiring them to write SQL.

• POC/exposure only; not represented as production ML model training or production Cortex deployment.

Healthcare Claims Data & Analytics Alignment

• Can design ETL/ELT patterns for healthcare datasets that combine claims, eligibility, provider, member/patient, and EHR-derived data while preserving lineage, validation, and privacy controls.

• Claims-focused data quality approach includes required-field validation, code-format checks, duplicate detection, source-to-target reconciliation, amount validation, rejected/denied claim monitoring, and audit-ready exception handling.

• Healthcare analytics use cases include denial trends, reimbursement analysis, utilization, payer performance, claim turnaround time, duplicate claims, and predictive/anomaly-detection opportunities.

EDUCATION & CERTIFICATIONS

Bachelor of Pharmacy (B.Pharm) - Gujarat University, India 2010

Snowflake SnowPro Core Microsoft Azure Data Engineer Associate (DP-203) Microsoft Power BI Data Analyst Associate (PL-300) Alteryx Designer Core UiPath RPA Developer Foundation Apache Airflow Fundamentals (Astronomer)



Contact this candidate