ARATI TARPARA
LEAD DATA ENGINEER AZURE DATABRICKS PYSPARK SNOWFLAKE DATA ARCHITECTURE
Allen, TX 75013 469-***-**** **************@*****.*** linkedin.com/in/arati-tarpara US Citizen
PROFESSIONAL SUMMARY
Lead Data Engineer with 15+ years of combined experience in enterprise data engineering, ETL/ELT, cloud platforms, database development, healthcare business systems, and business intelligence across insurance, healthcare, and manufacturing. Experienced leading large-scale insurance carrier data platforms using Snowflake, Python, Apache Airflow, Azure, ADF, ADLS Gen2, Nexla, and UiPath. Strong hands-on background in advanced SQL, stored procedures, reusable pipeline frameworks, data quality, reconciliation, production troubleshooting, CI/CD, and month-end processing. Led 400+ production pipelines supporting 200+ insurance carriers and helped reduce carrier processing timelines from 45 days to 14 days. Experience maps directly to modern Lakehouse and Medallion patterns, including raw/landing, standardized, cleansed, and business-ready data layers. Familiar with Azure Databricks, PySpark, Delta Lake, Synapse Analytics, and healthcare interoperability concepts including EHR data, FHIR, HL7, HIPAA-aware data handling, medical billing, coding, and insurance verification. Recent hands-on POC exposure includes Snowflake Cortex and Streamlit for Snowflake-native AI and natural-language analytics use cases.
TECHNICAL SKILLS
Cloud & Data Platforms: Snowflake, Microsoft Azure (Azure Data Factory, ADLS Gen2, Azure SQL, Synapse Analytics), Azure Databricks (working knowledge/exposure), AWS S3
Lakehouse & Big Data: PySpark/Spark concepts, Delta Lake, Lakehouse Architecture, Medallion Architecture, structured and unstructured data processing, scalable batch processing
Data Engineering & Orchestration: Apache Airflow (Astronomer), ETL/ELT, Nexla, Informatica PowerCenter 10.2, SSIS, UiPath RPA, Streamlit, Files.com
Programming & SQL: Python (Pandas, NumPy), SQL, T-SQL, PL/SQL, DDL/DML, stored procedures, functions, triggers, views, Bash, R, VBA
Snowflake & Data Quality: Snowflake SQL, stored procedures, layered transformations, audit logging, validation, reconciliation, duplicate controls, root-cause analysis, SLA monitoring
DevOps: Azure DevOps, Git, branch management, PR-triggered releases, DAG deployment, CI/CD, Dev-to-Prod promotion
Healthcare Data: EHR/EMR concepts, claims and billing data, FHIR/HL7 concepts, insurance verification, HIPAA-aware data handling, patient and clinical workflow data, medical claims lifecycle, professional/institutional claims concepts, ICD-10, CPT, HCPCS familiarity, eligibility, reimbursement, denials, payer/provider data, billed/allowed/paid amounts, deductible, copay, coinsurance
AI & Analytics Enablement: Snowflake Cortex (POC/exposure), Streamlit, LLM-enabled analytics concepts, Claude, Microsoft Copilot, Power BI, Tableau, predictive analytics concepts
Domain: Insurance carrier and commission data, policy/insured attributes, premium calculations, agent hierarchy, payment frequency, healthcare billing and insurance verification
PROFESSIONAL EXPERIENCE
Integrity Marketing Group LLC Dallas-Fort Worth, TX (Hybrid) Oct 2021 - Present
Senior Data Engineer Lead Data Engineer
• Lead end-to-end ingestion and processing of 10,000+ structured and unstructured files monthly from 200+ insurance carrier sources into Snowflake using Airflow, Python, Nexla, UiPath RPA, ADF, and custom processing modules.
• Directed architecture changes that reduced carrier data processing timelines from 45 days to 14 days, improving month-end reconciliation and SLA performance.
• Designed and supported reusable Snowflake transformation patterns using stored procedures, standardized tables and views, validation rules, duplicate controls, audit logging, and carrier-specific business logic.
• Own production troubleshooting and root-cause analysis across ingestion and transformation pipelines using logs, quarantine/error data, job-level tracking, reconciliation results, and Snowflake query analysis.
• Partner with business analysts, finance teams, and carrier stakeholders to translate commission feed requirements and source-to-target mappings into scalable technical solutions.
• Manage Azure DevOps CI/CD processes for DAGs and data pipelines, including branch management, code review, controlled releases, and Dev-to-Prod deployment.
• Mentor and guide distributed engineering teams, promoting reusable development standards, code quality, documentation, and production support practices.
• Led month-end reconciliation improvements and achieved full reconciliation within two weeks, including carrier mapping updates, baseline changes, exception resolution, and production validation.
• Built hands-on Snowflake Cortex and Streamlit POC work to explore natural-language interaction with governed Snowflake data and AI-enabled analytics; this work is exploratory rather than production ML model development.
• Designed pipelines around layered/Medallion data architecture principles, separating landing/raw ingestion, standardized processing, cleansed data, and business-ready outputs to improve traceability, reprocessing, and downstream reliability.
• Process both structured and unstructured source files and apply reusable Python-based parsing, standardization, schema validation, business-rule validation, metadata handling, and duplicate detection before downstream warehouse processing.
• Use Azure Data Factory and ADLS Gen2 for cloud ingestion and storage patterns and work with Azure Synapse concepts for analytics and warehousing use cases; experience is directly transferable to Databricks/Delta Lake Lakehouse implementations.
• Participate in technical design reviews and architecture decisions covering orchestration, storage, transformation, performance, data quality, failure handling, observability, scalability, and production support.
• Automate operational controls such as file-count reconciliation, missing-file detection, quarantine checks, failure logging, and Jira-based exception workflows to improve reliability and reduce manual monitoring.
• Apply data-quality controls throughout the pipeline including record/file count validation, required-column checks, duplicate controls, source-to-target reconciliation, exception quarantine, and audit logging before data reaches downstream consumers.
Lead Data Engineer ETL Pipeline Architect Oct 2021 - Oct 2024
• Architected a scalable, carrier-agnostic ETL and automation framework using Snowflake, Nexla, UiPath, Azure, and Airflow patterns, allowing new carrier feeds and formats to be onboarded through reusable components.
• Oversaw delivery of 400+ production pipelines with 1,000+ sub-pipeline branches supporting analytics, reconciliation, and operational reporting across 200+ insurance carriers in 44 states.
• Established ADF and ADLS Gen2 integrations for carrier file ingestion and reusable mapping documentation, strengthening data validation and reducing manual development effort.
• Led replacement of a third-party reconciliation service with an in-house UiPath + Snowflake solution, reducing external dependency and improving auditability and governance.
HOBI International Inc Dallas, TX Nov 2019 - Oct 2021
Data Analyst Contract
• Built Informatica PowerCenter 10.2 ETL pipelines to ingest Avro files from AWS S3, reconcile data with PostgreSQL, and deliver cleansed datasets to Snowflake for reporting.
• Designed Power BI dashboards for Sales, Production, and executive teams using DirectQuery, Live Connection, DAX measures, and calculated columns.
• Created and optimized PL/SQL stored procedures, functions, and packages and performed database performance tuning using indexing and query optimization techniques.
• Created GIS layers in ArcGIS Pro and ArcGIS Online and published spatial analytics to Power BI.
Mar-Lan Industries Irving, TX Jun 2013 - Nov 2019
Data Analyst Full-Time
• Designed SQL Server database objects including tables, stored procedures, views, triggers, and user-defined functions supporting OLAP and enterprise reporting.
• Developed and optimized SSIS/DTS ETL packages across heterogeneous source systems and authored complex T-SQL and MDX queries.
• Tuned SQL performance using execution plans, clustered/non-clustered indexes, and database monitoring; supported backup/restore, reconciliation, and database administration across 10+ environments.
Shivom USA Inc. Denton, TX Nov 2011 - Jun 2013
Business Consultant Self-Employed
• Supported hospital medical billing and coding processes and assisted with accurate patient billing workflows.
• Performed insurance verification activities to support reimbursement processing and reduce avoidable claim issues.
• Assisted with patient documentation, pharmacy technician duties, medication inventory, and healthcare administrative operations.
• Worked with healthcare operational data and workflows involving patient, insurance, billing, coding, and clinical documentation, providing foundational experience for modern EHR/claims-data engineering environments.
• Applied healthcare privacy and data-handling awareness when working with patient and insurance information; familiar with HIPAA expectations and healthcare data confidentiality requirements.
• Supported healthcare billing and insurance workflows involving patient, provider, payer, diagnosis, procedure, eligibility, and reimbursement information, giving direct exposure to the business context behind medical claims data.
• Familiar with the medical claims lifecycle from eligibility and claim submission through acceptance/rejection, adjudication, payment, and denial follow-up, including common claim attributes such as service date, diagnosis/procedure codes, billed amount, allowed amount, paid amount, deductible, copay, and coinsurance.
• Working familiarity with ICD-10, CPT, and HCPCS coding concepts and how coded clinical and billing information is used in claims processing, reimbursement, reporting, and healthcare analytics.
• Understand professional and institutional claims concepts, payer/provider relationships, duplicate/invalid claim checks, denial analysis, and HIPAA/PHI-aware handling of patient and insurance data.
SELECTED PROJECTS
Enterprise Carrier Data Platform - Snowflake + Airflow
• Directed a reusable dual-DAG orchestration pattern with dynamic task expansion and parallel Snowflake stored-procedure execution across STAGE and CLEANSED layers.
• Implemented transformation logic covering policy and insured attributes, status data, premium calculations, agent hierarchy, and payment frequency, with per-carrier/job audit logging.
• Designed automated file movement from carrier portals through Files.com and Streamlit into ADLS Gen2, followed by validation, duplicate checking, downstream Snowflake processing, and exception handling.
Azure Lakehouse / Databricks Alignment
• Current architecture uses Azure cloud storage, layered data processing, Python transformations, orchestration, data-quality controls, and CI/CD patterns that align closely with Databricks Lakehouse and Bronze/Silver/Gold design principles.
• Working knowledge of PySpark/Spark DataFrame processing, Delta Lake concepts, partitioning, scalable transformations, schema enforcement/evolution, and performance-oriented Lakehouse design; positioned as transferable experience rather than overstated production Databricks ownership.
Healthcare Data & Interoperability Exposure
• Healthcare domain background includes medical billing, coding, insurance verification, patient documentation, and healthcare administrative workflows.
• Familiar with modern healthcare integration concepts including EHR/EMR data, FHIR, HL7, claims/billing datasets, HIPAA-aware handling, and analytics use cases that combine clinical and operational data.
Snowflake Cortex / Streamlit AI POC
• Explored a Snowflake-native AI use case for natural-language access to governed carrier processing and reconciliation data.
• Focused on business context, reliable underlying data, user interaction through Streamlit, and how Cortex capabilities can support business users without requiring them to write SQL.
• POC/exposure only; not represented as production ML model training or production Cortex deployment.
Healthcare Claims Data & Analytics Alignment
• Can design ETL/ELT patterns for healthcare datasets that combine claims, eligibility, provider, member/patient, and EHR-derived data while preserving lineage, validation, and privacy controls.
• Claims-focused data quality approach includes required-field validation, code-format checks, duplicate detection, source-to-target reconciliation, amount validation, rejected/denied claim monitoring, and audit-ready exception handling.
• Healthcare analytics use cases include denial trends, reimbursement analysis, utilization, payer performance, claim turnaround time, duplicate claims, and predictive/anomaly-detection opportunities.
EDUCATION & CERTIFICATIONS
Bachelor of Pharmacy (B.Pharm) - Gujarat University, India 2010
Snowflake SnowPro Core Microsoft Azure Data Engineer Associate (DP-203) Microsoft Power BI Data Analyst Associate (PL-300) Alteryx Designer Core UiPath RPA Developer Foundation Apache Airflow Fundamentals (Astronomer)