Venkatesh Kala
Data Engineer / AI-ML Engineer
Hammond, IN 203-***-**** ****.*************@*****.*** LinkedIn
PROFESSIONAL SUMMARY
Data Engineer / AI-ML Engineer with 6+ years of experience building scalable Python, SQL, PySpark, and cloud-based data pipelines across financial, healthcare, and enterprise environments. Experienced with GCP, BigQuery, AWS, Snowflake, large-scale data processing, data quality, analytics, Tableau, and machine learning workflows. Delivers validated analytical datasets and ML-ready solutions that support reliable reporting and business insights.
TECHNICAL SKILLS
•Programming & Data Analysis: Python, SQL, Pandas, NumPy, Matplotlib, Seaborn, Jupyter Notebook, Exploratory Data Analysis (EDA)
•Data Engineering: ETL/ELT, Data Pipelines, PySpark, Apache Spark, Data Transformation, Data Integration, Data Cleansing, Data Validation, Data Profiling, Data Quality, Feature Engineering
•Machine Learning: Scikit-learn, Regression, Classification, Clustering, Time Series, Forecasting, Predictive Modeling, Anomaly Detection, Recommendation Systems, Feature Engineering, Model Evaluation, Hyperparameter Tuning
•Deep Learning & NLP: TensorFlow, PyTorch, Keras, NLP, Transformers, Text Classification, Named Entity Recognition, Embeddings, Semantic Search, Document Intelligence
•Generative AI: LLMs, Prompt Engineering, RAG, LangChain, LlamaIndex, OpenAI APIs, Hugging Face, Vector Databases, Embeddings, Semantic Search
•Big Data Technologies: Apache Spark, PySpark, Databricks, Distributed Data Processing, Batch Processing, Data Transformation
•Cloud & ML Platforms: AWS, Amazon S3, AWS Glue, SageMaker, Lambda, GCP, BigQuery, Google Cloud Storage, Vertex AI, Databricks, Azure, Azure Data Factory, Azure Data Lake, Azure Synapse, Dataproc, Cloud Composer
•Databases & Warehousing: Snowflake, PostgreSQL, MySQL, SQL Server, MongoDB, BigQuery
•Visualization: Power BI, Tableau, Excel
•MLOps & Deployment: MLflow, Docker, Kubernetes, Git, GitHub Actions, CI/CD, REST APIs, Flask, FastAPI, Model Monitoring
•Methodologies & Tools: Jira, Agile/Scrum, SDLC, VS Code, GitHub
CERTIFICATIONS
•Hugging Face Agents Course – Certificate of Excellence
•Fundamentals of Machine Learning and Artificial Intelligence certification
•Claude with the Anthropic API certification
•Google Agile Essentials Certificate
EDUCATION
New England College
Aug 2022 - Dec 2023
Master's, Computer Information
•GPA: 4.0
Henniker, NH USA
Osmania University
Jun 2016 - Apr 2019
Bachelor's
Hyderabad, India
PROFESSIONAL EXPERIENCE
Chase Jul 2025 - Present
Data Engineer / AI-ML Engineer Indianapolis, IN
Data Engineer / AI-ML Engineer supporting financial data pipelines, analytics, and machine learning workloads. I build and optimize Python, SQL, PySpark, Databricks, AWS, and Snowflake solutions for high-volume customer and transaction data, with a focus on ETL/ELT, data quality, validation, reconciliation, and ML-ready datasets.
•Environment: Python, SQL, Pandas, NumPy, PySpark, Apache Spark, Scikit-learn, TensorFlow, PyTorch, Databricks, AWS S3, AWS Glue, SageMaker, Lambda, Snowflake, MLflow, Docker, FastAPI, Flask, OpenAI APIs, LangChain, RAG, Vector Databases, Power BI, Tableau, Git, Jira, Agile/Scrum.
•Design and develop Python and SQL-based data pipelines to ingest, transform, validate, and prepare large-scale financial, customer, transaction, and operational datasets for analytics and machine learning applications.
•Develop scalable ETL/ELT workflows using Python, PySpark, Apache Spark, and cloud-based services to process structured and semi-structured enterprise data.
•Performed data profiling, cleansing, validation, transformation, reconciliation, and quality checks to improve reliability and consistency of downstream analytical datasets.
•Built reusable PySpark transformation frameworks for filtering, joining, aggregating, deduplicating, and transforming high-volume datasets.
•Developed feature engineering pipelines using Python, SQL, Pandas, and PySpark to prepare model-ready datasets for predictive
analytics and machine learning applications.
•Built classification, regression, clustering, anomaly detection, and predictive modeling solutions using Scikit-learn, TensorFlow, and
PyTorch.
•Developed analytical models to identify customer behavior patterns, transaction trends, operational anomalies, and other
business-relevant patterns within financial datasets.
•Implemented data validation and reconciliation processes to compare source and target datasets and identify missing, duplicate,
inconsistent, or invalid records.
•Used AWS services including Amazon S3, AWS Glue, SageMaker, and Lambda to support data ingestion, transformation, storage,
machine learning experimentation, and deployment workflows.
•Worked with Snowflake to develop analytical datasets, optimize SQL queries, create feature datasets, and support downstream
reporting and ML workloads.
•Used Databricks and Apache Spark for large-scale data processing, transformation, feature engineering, and distributed analytical
workloads.
•Implemented data quality checks and exception-handling workflows to identify pipeline failures, schema changes, missing data, and
data integrity issues.
•Developed SQL queries involving complex joins, subqueries, aggregations, window functions, CTEs, and analytical transformations.
•Optimize SQL and PySpark workloads to improve processing efficiency and support large-scale data pipeline execution.
•Supported MLOps workflows using MLflow, Docker, Git, CI/CD pipelines, model tracking, and model monitoring practices.
American Board of Internal Medicine (ABIM)
Jan 2024 - Jun 2025
Data Engineer / Machine Learning Engineer
Philadelphia, PA
Data Engineer / Machine Learning Engineer supporting healthcare, examination, and operational data initiatives. I develop Python, SQL, PySpark, AWS, GCP, BigQuery, and Snowflake pipelines; build validated analytical datasets; and apply machine learning and NLP
techniques for reporting, analytics, and data-quality use cases.
•Environment: Python, SQL, Pandas, NumPy, PySpark, Apache Spark, Scikit-learn, TensorFlow, PyTorch, NLP, Transformers, AWS, GCP,
BigQuery, Dataproc, Cloud Composer, Snowflake, Databricks, Power BI, Tableau, OpenAI APIs, LangChain, RAG, Vector Databases, Git,
Jira, Agile/Scrum.
•Developed Python and SQL-based data pipelines to process healthcare, examination, assessment, operational, and enterprise
datasets.
•Designed ETL/ELT workflows to extract, cleanse, transform, validate, and load structured and semi-structured data from multiple
source systems.
•Used Pandas, PySpark, and SQL to transform raw healthcare datasets into standardized analytical and machine learning datasets.
•Performed data profiling, exploratory data analysis (EDA), statistical analysis, and data-quality assessments across large
healthcare-related datasets.
•Implemented data cleansing and validation processes to identify missing values, duplicate records, inconsistent formats, invalid
values, and source-system data issues.
•Developed reusable PySpark pipelines for large-scale data transformation, aggregation, filtering, joining, and feature engineering.
•Built machine learning models for classification, prediction, segmentation, performance analytics, and anomaly detection using
Scikit-learn and TensorFlow.
•Created feature engineering workflows using Python, SQL, Pandas, and PySpark to prepare model-ready datasets.
•Applied NLP techniques to text-based healthcare and examination datasets including tokenization, text preprocessing, text
classification, embeddings, and named entity recognition.
•Worked with transformer-based models for text analytics, document processing, semantic search, and intelligent information
retrieval.
•Developed anomaly-detection and data-quality models to identify unusual records, inconsistent values, and potential data integrity
issues.
•Built reusable preprocessing pipelines for missing-value treatment, duplicate handling, categorical transformations, normalization,
validation, and feature selection.
•Worked with Dataproc, Cloud Composer, and other cloud storage and compute services to support scalable data processing, analytical
workloads, model experimentation, and deployment.
•Worked with AWS and GCP services to support cloud-based data processing, storage, analytical workloads, and machine learning
workflows.
•Used BigQuery and Snowflake for analytical data processing, SQL-based transformations, reporting datasets, and downstream ML
workloads.
•Developed data pipelines supporting reporting and analytical dashboards and provided validated datasets to downstream reporting
teams.
•Developed and optimized Power BI and Tableau dashboards to present analytical findings, operational metrics, data-quality trends,
and model results to stakeholders.
DXC Technology
Jan 2019 - Jun 2022
Data Engineer / ML Analyst
India
Data Engineer / ML Analyst supporting enterprise business and operational data processing. I used Python, SQL, Pandas, and relational databases to build ETL workflows, automate data preparation and validation, develop analytical datasets, and support predictive analytics, reporting, reconciliation, and technical documentation.
•Environment: Python, SQL, Pandas, NumPy, Scikit-learn, Excel, Power BI, MySQL, SQL Server, ETL, Data Validation, Data Analysis, Data Pipelines, Machine Learning, Git, JIRA, Agile/Scrum.
•Developed Python and SQL-based data processing workflows to extract, transform, cleanse, validate, and analyze enterprise business and operational datasets.
•Supported ETL activities by extracting data from relational sources, applying transformation rules, validating source-to-target mappings, and preparing downstream analytical datasets.
•Developed SQL queries to extract, join, transform, aggregate, and analyze data from multiple relational database sources.
•Used Python, Pandas, and NumPy to automate data preparation, cleansing, transformation, validation, and recurring analytical tasks.
•Performed exploratory data analysis and statistical analysis to identify trends, patterns, anomalies, and performance drivers within business datasets.
•Built data preparation and feature engineering workflows using Python, Pandas, NumPy, and SQL for predictive analytics and machine learning applications.
•Developed predictive models using Scikit-learn for classification, regression, forecasting, clustering, and anomaly detection use cases.
•Worked with time-series datasets to identify trends, seasonality, operational patterns, and forecasting opportunities.
•Performed data cleansing and validation to handle missing values, duplicate records, outliers, inconsistent records, and source-system issues.
•Supported development of scalable data pipelines and analytical workflows for business reporting and machine learning use cases.
•Developed reusable Python scripts to automate recurring data preparation, data validation, reporting, and reconciliation activities.
•Supported data integration activities by reviewing source data, transformation logic, target datasets, and source-to-target reconciliation results.
•Developed SQL-based analytical datasets to support reporting, business analysis, and downstream data science activities.
•Performed data-quality checks and investigated discrepancies between source and transformed datasets.
•Maintained technical documentation for SQL queries, data definitions, data mappings, analytical workflows, reporting logic, and model-development processes.