YASHASHREE REDDY KARRI
312-***-**** • ********************@*****.***
PROFESSIONAL SUMMARY:
Data Engineer and AI/ML Engineer with 5 years of experience designing scalable data platforms, building end-to-end ETL/ELT pipelines, and delivering AI-driven solutions across enterprise environments. Expertise in Python, SQL, Apache Spark, Airflow, AWS, Databricks, and modern Generative AI technologies including LLMs, LangChain, LangGraph, RAG, and multi-agent systems. Proven track record of developing production-ready machine learning models, optimizing large-scale data processing workflows, and transforming complex datasets into actionable business insights. Skilled in data engineering, advanced analytics, MLOps, and cloud-native architectures, with strong experience collaborating with cross-functional teams to deliver secure, scalable, and high-impact data and AI solutions. EDUCATION:
Illinois Institute of Technology, Chicago, IL
Master of Science in Computer Science (2025)
GITAM Deemed to be University, Visakhapatnam, India Bachelor of Engineering in Computer Science and Business Systems (2023) TECHNICAL SKILLS:
Programming & Query Languages: Python, SQL, C, C++ Data Engineering: Apache Spark, Apache Airflow, ETL/ELT Pipelines, Data Modeling, Data Warehousing, Data Lakes, Batch & Streaming Pipelines
Generative AI : Machine Learning (ML),Deep Learning (DL), NLP, LLMs, LangChain, Open AI API, Hugging face, TensorFlow, PyTorch, Scikit-Learn, ChromaDB, RAG, AgenticAI, Vector Databases, Prompt Engineering, Multi-Agent Systems Databases: PostgreSQL, SQL Server, NoSQL (concepts) Data Science & Analytics: EDA, Statistical Analysis, Feature Engineering, Hypothesis Testing, Predictive Modeling Machine Learning: Supervised & Unsupervised Learning, Model Evaluation, Hyperparameter Tuning Cloud & Big Data: AWS (S3, EC2, Lambda, SageMaker), Databricks Visualization: Tableau, Excel, Dashboarding, KPI Reporting Tools: Docker, Git, Jupyter, Anaconda, MLflow
WORK EXPERIENCE
Gainsight, USA Data Engineer / Data Scientist (AIML) July 2024 – Present
Roles & Responsibilities:
● Designed and implemented end-to-end ETL pipelines to ingest data from APIs, Excel, and SaaS platforms into Gainsight NXT.
● Proficient in Python, SQL, PyTorch, and Scikit-Learn with hands-on experience in Transformers, multi-agent frameworks, and NLP algorithms; well-versed in Git, CI/CD pipelines, and standard version control workflows.
● Built and deployed end-to-end AI agents leveraging LLMs, RAG, Google ADK, and LangGraph, leading large-scale agentic framework projects from ideation to production with measurable impact.
● Performed data cleaning, transformation, and normalization to ensure high-quality datasets for analytics.
● Proficient in AI-assisted coding tools (Claude Code, Copilot, Cursor) as a daily driver, with hands-on experience demonstrated through active code commits and repositories.
● Possess strong transferable experience in data engineering, analytics, and machine learning, with expertise in building secure, scalable data solutions and deriving actionable insights applicable to banking and financial services.
● Developed and optimized dashboards and reports to analyze customer health scores, engagement, and adoption metrics. Built scalable data models to support customer analytics, churn prediction, and lifecycle tracking.
● Applied data analysis techniques to identify trends, anomalies, and business insights.
● Automated workflows using rules engines to trigger actions based on real-time customer data.
● Implemented data validation and monitoring frameworks to maintain data integrity and reliability.
● Optimized performance of large-scale data processing workflows handling high-volume customer datasets.
● Collaborated with stakeholders to translate business requirements into data-driven solutions.
● Experience in financial markets (delta one, store of value, and/or FICC options trading), plus Linux-based, concurrent, high-throughput, low-latency systems; functional programming experience is a plus.
● Worked with semi-structured data and performed transformations for reporting and ML use cases. Smartknowers, USA Data Scientist / Data Analyst
Aug 2023 – July 2024
Roles & Responsibilities:
● Performed exploratory data analysis (EDA) to uncover patterns, correlations, and trends in datasets.
● Built and deployed machine learning models for classification and predictive analytics.
● Designed data preprocessing pipelines including feature engineering, encoding, normalization, and handling missing values.
● Applied statistical techniques and hypothesis testing to validate data insights.
● Conducted model evaluation using metrics such as accuracy, precision, recall, F1-score, and ROC-AUC.
● Performed hyperparameter tuning and cross-validation to improve model performance.
● Created visualizations and dashboards to communicate insights to stakeholders.
● Worked with large datasets to derive actionable insights and support business decisions.
● Automated data workflows and analysis pipelines using Python.
● Assisted in deploying models using APIs for real-time inference.
● Documented analytical processes, models, and insights for business and technical teams. Sonata, India Data Scientist / Machine Learning Engineer Mar 2022 – Apr 2023
Roles & Responsibilities:
● Built deep learning models (CNN, RNN, GRU) for predictive maintenance and anomaly detection.
● Developed a Conditional GAN for data augmentation, improving recall and model performance on rare events.
● Designed time-series forecasting models to predict equipment health and remaining useful life.
● Built and maintained data pipelines for ingesting and processing sensor data streams.
● Implemented real-time data processing for anomaly detection and alerting systems.
● Developed MLOps pipelines using AWS (SageMaker, Lambda) for automated training and deployment.
● Managed model versioning, monitoring, and retraining using MLflow.
● Integrated model outputs into downstream systems for decision-making and automation.
● Optimized inference latency for real-time applications.
● Collaborated with engineering teams to deploy scalable AI solutions in production.
● Applied GenAI concepts by integrating AI outputs into chatbot-based interfaces for engineers. PROJECTS
Multi-Agent GenAI Enterprise Automation System
● Built a multi-agent system using LangChain, LangGraph, and RAG for enterprise workflow automation.
● Designed data pipelines for document ingestion, embedding generation, and semantic retrieval.
● Integrated vector databases for efficient information retrieval.
● Automated generation of BRDs, data mappings, and test cases using LLMs.
● Applied prompt engineering to improve response quality and consistency.
● Developed a dashboard for role-based interaction and workflow execution. PeerTherapy – AI-Driven Platform
● Designed backend systems for real-time chat data processing and storage.
● Applied NLP techniques for sentiment analysis and user support recommendations.
● Built analytics pipelines to track engagement and user behavior.
● Ensured secure and privacy-focused data handling. GPS-Based Vehicle Theft Detection System
● Built real-time data ingestion pipelines using GPS-based tracking.
● Processed streaming data for anomaly detection and alert generation.
● Designed dashboards for monitoring vehicle activity and alerts.
● Applied data analytics to improve detection accuracy and response time. ADDITIONAL CAPABILITIES
● Strong understanding of data lifecycle: ingestion, transformation, storage, and visualization.
● Experience with distributed data systems and big data architectures.
● Hands-on experience with end-to-end ML pipelines and deployment.
● Knowledge of data governance, validation, and quality frameworks.
● Ability to solve complex business problems using data-driven approaches.
● Experience integrating GenAI solutions into enterprise workflows. KEY HIGHLIGHTS
● Strong combination of Data Engineering, Data Science, and Data Analytics expertise.
● Experience building scalable data systems and AI-driven applications.
● Proven ability to deliver insights and automation using modern data and GenAI technologies.