Menlo Park, CA
Phone: 415-***-****
Email: *********@*****.***
LinkedIn US Citizen Parisa Pazooki
Data Analytics Data Management Data Automation EDUCATION
M.Sc. in Biotechnology (Focus: Bioinformatics) University of San Francisco (USF) San Francisco, CA August 2020 – December 2022 M.Sc. in Genetics Azad University, North Tehran Branch Tehran, Iran September 2014 – September 2017 B.Sc. in Cellular and Molecular Biology Azad University, North Tehran Branch Tehran, Iran September 2006 – September 2010 TECHNICAL SKILLS
• Data Analysis & Programming: SQL, Python, Bash; pandas, NumPy; data cleaning, transformation, validation, integration
• Databases & SQL: Google BigQuery, PostgreSQL; relational databases, tables and schemas, SQL querying, joins, aggregations, filtering; DML (INSERT, UPDATE)
• Cloud & Data Workflows: Google Cloud Platform (GCP), AWS (S3, EC2, Batch); automated data processing, ETL workflows, workflow monitoring, troubleshooting
• Development & Collaboration: Linux/UNIX, Docker, Git, Jira, Confluence; technical documentation, issue tracking PROFESSIONAL EXPERIENCE
Bioinformatics Scientist Pragma Biosciences
South San Francisco, CA Apr 2024 — Present
• Query and maintain large biological and clinical datasets in Google BigQueryusingSQL, working with relational tables and schemas, joins, filtering, aggregations, and DML operations.
• Develop Python and SQL data workflows to integrate datasets from multiple sources, create and maintain analysis-ready tables, and validate schema consistency, missing values, duplicates, and data integrity.
• Monitor data-processing workflows, troubleshoot data and pipeline issues, and evaluate potential downstream impacts on analytical results.
• Collaborate with cross-functional teams to investigate data issues, document workflows, and communicate findings using Jira and Confluence.
Bioinformatics Scientist PawCo Foods
San Francisco, CA Jan 2023 — Apr 2024
• Built and maintained Python, SQL, and Bash data pipelines to query, transform, and integrate structured datasets across relational tables, using joins, aggregations, and DML operations.
• Developed reusable Python and SQL workflows to prepare and maintain structured datasets, validate table-level data quality, and ensure consistency across downstream data processes.
• Monitored pipeline execution, investigated processing failures and data-quality issues, and resolved problems affecting downstream outputs.
• Optimized data-processing workflows, reducing computational costs by more than 40% while improving efficiency and reproducibility.
Research Intern Recombia Biosciences
South San Francisco, CA Jun 2022 — Sep 2022
• Developed Python-based tools to retrieve, process, and organize biological data from public databases, including UniProt.
• Automated data-processing and validation tasks to improve data consistency and reduce manual processing.
• Analyzed and organized sequence and experimental datasets to support downstream research and reporting. Research Intern University of California San Francisco San Francisco, CA Jun 2021 — Dec 2021
• Processed, cleaned, and analyzed genomic datasets using Python, R, and command-line tools.
• Performed data quality checks and maintained reproducible workflows for downstream analysis.
• Integrated data from multiple sources and prepared structured outputs, visualizations, and summaries for research teams. Research Scientist Pasteur Institute of Iran
Tehran, Iran Oct 2017 — Nov 2019
• Analyzed genetic and clinical data using chi-squared and Fisher’s exact tests to investigate factors associated with chronic myeloid leukemia.
• Evaluated genotype–phenotype associations and integrated clinical variables to support cancer-focused biomedical research. SELECTED PUBLICATIONS
Key genes and regulatory networks involved in the initiation, progression, and invasion of colorectal cancer; Future Science 2018