Post Job Free
Sign in

Senior AI Python Data Engineer

Location:
North Brunswick, NJ
Posted:
October 08, 2026

Contact this candidate

Resume:

KARTHIK KAKARLA

Senior AI Python & Data Engineer

Python Application Development Advanced SQL Financial Markets Portfolio & Market Data Systems AI-Assisted Engineering

*****************@*****.*** 218-***-**** LinkedIn: www.linkedin.com/in/karthik-kakarla-841336416

PROFESSIONAL SUMMARY

Senior AI Python & Data Engineer with 11+ years of experience delivering enterprise Python, data, AI/ML, and production engineering solutions across financial services, healthcare, retail.

Strong hands-on Python engineering background spanning backend services, REST APIs, FastAPI/Flask, data processing, object-oriented development, debugging, testing, code reviews, and production support.

Advanced SQL and relational database engineering experience including complex queries, stored procedures, views, data modeling, indexing, partitioning, query optimization, and performance tuning for high-volume datasets.

Financial-services experience supporting investment-oriented applications and data workflows, with hands-on understanding of equities, fixed income, securities and financial instruments, trade lifecycle, orders and executions, positions and holdings, transactions, pricing and valuation, and market data.

Built and supported portfolio management and real-time market data capabilities, combining Python services, SQL/data platforms, APIs, streaming pipelines, and governed cloud infrastructure for investment use cases.

Experienced collaborating with business stakeholders and technical teams to translate requirements into scalable application and data solutions, with ownership across design, implementation, testing, deployment, monitoring, and troubleshooting.

Hands-on with AI-assisted development tools including GitHub Copilot and Claude for coding, debugging, testing, code review, documentation, refactoring, and developer productivity, alongside production Generative AI engineering.

Cloud and software delivery experience across Azure and AWS with Docker, Kubernetes, Terraform, GitHub Actions/Jenkins/Azure DevOps, automated CI/CD, observability, and engineering standards.

TECHNICAL SKILLS

TECHNICAL CATEGORY

SKILLS / TECHNOLOGIES

Programming & Python Engineering

Python, SQL, Java, R, JavaScript, TypeScript, Scala; OOP, asynchronous Python, Pandas, NumPy, pytest/unit testing, debugging, code review

Backend & APIs

FastAPI, Flask, Django, REST APIs, Microservices, JSON, enterprise application integration

Financial Markets & Investment Systems

Equities, Fixed Income, Securities & Financial Instruments, Trade Lifecycle, Orders & Executions, Positions & Holdings, Transactions, Pricing & Valuation, Market Data, Portfolio Management Systems, Real-Time Market Data Applications

Databases & Data Modeling

Oracle, SQL Server, Snowflake, PostgreSQL, MySQL, MongoDB, Redshift, Teradata; Complex SQL, Stored Procedures, Views, Relational Data Modeling, Query Optimization, Performance Tuning

Data Engineering

PySpark, Apache Spark, Databricks, Apache Kafka, Airflow, dbt, Delta Lake, AWS Glue, Azure Data Factory, Hive, Hadoop, Denodo, ETL/ELT

AI & AI-Assisted Development

GitHub Copilot, Claude, LLMs, RAG, LangChain, LangGraph, OpenAI GPT, Prompt Engineering, Hugging Face Transformers, AI Agents

Cloud / DevOps / MLOps

Azure, AWS, GCP, Docker, Kubernetes, Terraform, MLflow, GitHub Actions, Jenkins, Azure DevOps, CI/CD, Azure Monitor, Application Insights, CloudWatch

Operating Systems

Windows, Linux, Mac OS, Unix Shell Scripting

PROFESSIONAL EXPERIENCE

Edward Jones – St. Louis, Missouri

Oct 2024 – Present

Senior Generative AI Data Engineer

Designed and developed Python-based enterprise services and APIs supporting investment research, portfolio analysis, advisor workflows, and financial data consumption, translating business requirements into maintainable production applications.

Built and supported portfolio management capabilities using Python, SQL, REST APIs, Snowflake, and Azure services to process positions, holdings, transactions, pricing, valuation, and market data for investment-oriented workflows.

Engineered real-time market and client data pipelines with Python/PySpark, Databricks, Apache Kafka, and Delta Lake, enabling timely downstream portfolio analytics, investment research, and advisor-facing applications.

Developed application logic and data workflows around financial-market concepts including equities, fixed income, securities, orders and executions, trade lifecycle events, positions, holdings, and transaction data.

Created and optimized complex SQL queries, transformations, and analytical data models for high-volume financial datasets, applying query tuning and data-quality validation to support reliable portfolio and market-data processing.

Built high-throughput FastAPI services and microservice integrations for financial applications, implementing validation, exception handling, logging, and secure REST interfaces for downstream consumers.

Architected an enterprise Generative AI platform using Python, Azure OpenAI, and LangChain to assist financial advisors with investment research and recommendations, supporting more than 3,000 advisors and reducing manual research time by 40%.

Designed a RAG framework with LangChain, Pinecone, and Azure AI Search over more than 2M investment documents, grounding responses in approved financial content and improving answer relevance by 35%.

Developed multi-agent portfolio-analysis workflows with LangGraph and AI agents, orchestrating tool calls and data access for complex investment research tasks.

Applied GitHub Copilot and Claude as AI-assisted engineering tools for Python development, debugging, unit-test generation, code review, refactoring, technical documentation, and routine code-maintenance activities.

Implemented automated testing and participated in peer code reviews, debugging, release validation, and production support for Python, API, data, and AI components, resolving application and pipeline issues through root-cause analysis.

Automated deployment and infrastructure workflows using CI/CD, Docker, Kubernetes, Terraform, GitHub Actions, and MLflow, supporting repeatable releases and cloud-native application delivery.

Implemented observability and operational controls with Azure Monitor, Application Insights, MLflow, and structured logging to monitor application, model, retrieval, and data-pipeline behavior in production.

Collaborated with investment-domain stakeholders, engineering teams, architects, and data teams to refine requirements, review technical designs, improve engineering standards, and own solutions from development through production support.

Environment: Python, SQL, FastAPI, REST APIs, PySpark, Databricks, Apache Kafka, Snowflake, Azure OpenAI, LangChain, LangGraph, RAG, Pinecone, Azure AI Search, MLflow, GitHub Copilot, Claude, Docker, Kubernetes, Terraform, GitHub Actions, Azure Cloud, Power BI

Northwell Health – New Hyde Park, NY

May 2022 – Sept 2024

Senior Data Engineer AI & Machine Learning

Architected Python and PySpark data-processing applications on Databricks for large-scale clinical, EHR, lab, and claims datasets, with end-to-end ownership from requirements through production support.

Built scalable ETL/ELT frameworks with Apache Airflow, Azure Data Factory, dbt, and Delta Lake, integrating enterprise healthcare sources and reusable SQL transformations.

Developed and optimized complex SQL transformations, data models, validation logic, and curated datasets for downstream analytics and operational reporting.

Created secure FastAPI and Flask services exposing governed datasets to downstream applications and teams through REST-based interfaces.

Implemented real-time ingestion and monitoring pipelines using Apache Kafka, PySpark Streaming, and Azure Event Hubs for continuous high-volume data processing.

Built predictive and NLP solutions using Scikit-learn, XGBoost, TensorFlow, PyTorch, SpaCy, NLTK, and Hugging Face Transformers.

Automated testing, data-quality validation, deployment, and release workflows with Python, SQL, MLflow, Docker, and Azure DevOps CI/CD.

Established production monitoring with Azure Monitor, Application Insights, and MLflow and performed troubleshooting and root-cause analysis for data and model services.

Collaborated with clinicians, researchers, data scientists, application teams, and architecture stakeholders on technical design, code reviews, delivery standards, and production implementation.

Environment: Python, SQL, PySpark, Databricks, Apache Spark, Apache Airflow, dbt, Azure Data Factory, Apache Kafka, Azure Event Hubs, Delta Lake, Snowflake, Scikit-learn, TensorFlow, PyTorch, MLflow, FastAPI, Flask, Docker, Azure DevOps, FHIR, HL7, Azure Cloud

Pacific Western Bank – Beverly Hills, CA

Aug 2019 – May 2022

Senior Data & ML Engineer

Developed Python, PySpark, and SQL solutions for high-volume commercial lending, transaction, risk, and customer datasets within a financial-services environment.

Built scalable ETL/ELT workflows with Apache Airflow, AWS S3, Hive, and AWS Glue, integrating core banking sources and downstream analytics platforms.

Developed hands-on AWS Glue ETL jobs and Glue Data Catalog integrations using Python and PySpark to ingest, transform, and catalog banking datasets in Amazon S3.

Engineered real-time transaction ingestion using Apache Kafka, Spark Streaming, and AWS Lambda for continuous processing and monitoring.

Developed fraud-detection, credit-risk, segmentation, and portfolio analytics models using Python, Scikit-learn, XGBoost, and Random Forest techniques.

Created RESTful Python services with Flask and FastAPI to expose banking analytics and data capabilities to enterprise applications.

Optimized PySpark and SQL processing through partitioning, transformation tuning, and data-model improvements for large-scale financial datasets.

Implemented ML deployment and software-delivery workflows using MLflow, Docker, Jenkins CI/CD, testing, logging, and production monitoring.

Used Denodo for data virtualization across relational and analytical sources, building reusable virtual views with joins, filters, and aggregations.

Collaborated with risk analysts, data scientists, and engineering teams to deliver financial data, machine-learning, and cloud solutions through production support.

Environment: Python, SQL, PySpark, Apache Spark, Apache Airflow, Apache Kafka, Spark Streaming, AWS S3, AWS Glue, AWS Lambda, Hive, Scikit-learn, XGBoost, MLflow, FastAPI, Flask, Docker, Jenkins, Snowflake, Denodo, AWS Cloud

Carter’s, Inc. – Atlanta, Georgia

Jan 2017 – July 2019

Senior Python Data Engineer

Designed and developed Python, PySpark, SQL, and Apache Spark data pipelines for retail sales, customer, inventory, e-commerce, and in-store datasets.

Built ETL/ELT and workflow orchestration with Apache Airflow, Hive, HDFS, Python, and shell scripting, supporting reliable daily data operations.

Engineered real-time ingestion with Apache Kafka and Spark Streaming and optimized large-scale transformations through SQL and partitioning strategies.

Built centralized AWS S3/Hive/HDFS data-lake solutions, automated data-quality checks, and supported cloud migration and production troubleshooting.

Created customer segmentation and loyalty datasets plus Tableau/Power BI reporting for sales, inventory, and customer analytics.

Environment: Python, SQL, PySpark, Apache Spark, Apache Airflow, Apache Kafka, Spark Streaming, Hive, HDFS, AWS S3, AWS EC2, Flask, REST APIs, Apache Atlas, Tableau, Power BI, AWS Cloud

CitiusTech – Mumbai, Maharashtra

Aug 2015 – Dec 2016

Python Data Engineer

Developed Python, SQL, Pandas, NumPy, and Unix shell solutions for healthcare and clinical data processing and enterprise ETL workloads.

Built automated ETL workflows with Python, Informatica, Oracle, and SQL Server, integrating multiple healthcare applications and HL7 sources.

Designed and optimized complex SQL queries, stored procedures, views, indexing, and partitioning strategies for reporting, analytics, and database performance.

Built batch-processing frameworks and enterprise scheduling workflows to meet daily reporting SLAs and support production operations.

Created dimensional data models using star/snowflake schemas and data-warehousing practices and supported data migration and release management.

Provided production support through incident management, monitoring, troubleshooting, root-cause analysis, and application/data-pipeline issue resolution.

Environment: Python, SQL, Pandas, NumPy, Informatica, Oracle, SQL Server, HL7, Unix Shell Scripting, Data Warehousing, Star Schema, Snowflake Schema, Git, SVN, Data Modeling

EDUCATION

Cambridge Institute of Technology – Bengaluru, India

Aug 2011 – June 2015

Electronics and Communication Engineering (ECE)



Contact this candidate