Post Job Free
Sign in

Data Entry Specialist & PDF Automation Intern

Location:
Mumbai, Maharashtra, India
Salary:
3000 dollars
Posted:
October 06, 2026

Contact this candidate

Resume:

DARSH TUNGARIA

Data Entry Specialist Remote Internship Candidate (US)

Mumbai, India +91-702**-***** **************@*****.*** PROFESSIONAL SUMMARY

Detail-oriented data entry professional with two data entry internships converting PDF product and business documents into web-based systems with a strong focus on accuracy and turnaround. Skilled in Excel, Google Workspace and PDF tools, and able to build Python automation that cuts repetitive typing. Reliable, confidential and comfortable working remotely, with availability that can be aligned to US business hours. CORE SKILLS

Data Entry: Typing speed of 35 WPM, accuracy-first entry, record verification, duplicate and error checks, data cleaning, PDF-to-web-form entry, product catalog and pricing data Software: Microsoft Excel, Word and PowerPoint; Google Sheets, Docs and Forms; Adobe Acrobat and PDF viewers; Gmail and Outlook; Chrome-based web portals

Automation: Python (pdfplumber, Playwright), browser scripting, workflow automation for repetitive entry Professional: Time management, attention to detail, data confidentiality, clear written English PROFESSIONAL EXPERIENCE

Data Entry Intern — Prism Data Solutions

1-month internship

• Entered and verified product records (names, style codes, pricing, materials, sizes and descriptions) from PDF documents into online systems.

• Reviewed every entry against the source file before submission to keep error rates minimal.

• Flagged missing or inconsistent data to supervisors and resolved discrepancies quickly.

• Met daily targets and deadlines while maintaining consistent quality. Data Entry Intern — Winsoftec

1-month internship

• Maintained and updated structured records in spreadsheets and web portals with high accuracy.

• Cleaned and standardized data formats to improve consistency across records.

• Followed company procedures for data security and handled sensitive information with discretion. PROJECTS

PDF-to-Web-Form Automation Tool — Python, pdfplumber, Playwright

• Built a script that reads product PDFs, extracts 37 labeled fields (including multi-line descriptions) and fills the web data-entry form automatically.

• Skips N/A values, reports any field it cannot fill, and leaves final review to the operator to prevent bad submissions.

• Full source code attached in the appendix (pages 2–3). EDUCATION

Secondary School (Class X) — OLPS High School, Mumbai, India CGPA: 9.1 / 10

LANGUAGES

English — Fluent (spoken and written)

APPENDIX — PROJECT SOURCE CODE

PDF-to-Web-Form Automation Tool (fill_forms.py). Reads each PDF in a folder, extracts the labeled fields and fills the matching form boxes in the browser.

"""

Reads every PDF in the 'pdfs' folder and fills your data-entry form. Setup (Chromebook Linux terminal):

pip install pdfplumber playwright

playwright install chromium

Run:

python3 fill_forms.py

"""

import re

import sys

from pathlib import Path

import pdfplumber

from playwright.sync_api import sync_playwright

# SETTINGS (edit these) SITE_URL = "https://YOUR-SITE-ADDRESS-HERE" # page where you log in PDF_FOLDER = Path("pdfs") # put your PDFs here

SKIP_NA = True # leave fields blank when the PDF says N/A SUBMIT_AUTOMATICALLY = False # False = you review, then click submit yourself SUBMIT_BUTTON_TEXT = "Submit" # text on your site's submit button

# If a form label differs from the PDF label, add it here:

# "PDF label": "Form label"

FIELD_MAP = {

# "Fabric/Materials": "Fabric",

}

# LABELS = [

"Product Name", "Men Style Code", "Ladies Style Code", "Kids Style Code",

"Fabric/Materials", "Brand", "Manufacturer", "Product Price",

"Men Product Price", "Ladies Product Price", "Kids Product Price",

"Product Features 1", "Product Features 2", "No. of Colors Available",

"Product", "Country of Origin", "Generic Name", "Department", "Packers",

"Item Weight", "No. Of Sizes Available", "Length", "Care Instruction",

"Occasion type", "Product Dimensions", "Net Quantity", "Available Stock",

"Manufacturing Date", "Minimum Order Qty", "Maximum Order Qty",

"Product Details", "Item Number", "Discount In %", "Neckline",

"Sleeve Style", "Returns Days Policy", "Payment Mode",

]

LABEL_RE = re.compile(

r + join(re.escape(l) for l in LABELS) + r")\s*:\s* $", re.IGNORECASE,

)

CANON = {l.lower : l for l in LABELS}

def parse_pdf(path):

"""Return {label: value}. Multi-line values (Product Details) are joined text = ""

with pdfplumber.open(path) as pdf:

for page in pdf.pages:

text += (page.extract_text or "") + "\n"

data, current = {}, None

for line in text.splitlines :

line = line.strip if not line:

continue

m = LABEL_RE.match(line)

if m:

current = CANON[m.group(1).lower ]

data[current] = m.group(2).strip elif current:

data[current] += " " + line

return data

def fill_field(page, label, value):

form_label = FIELD_MAP.get(label, label)

# Find the first input/textarea/select after the label text. xpath = (

"xpath not ][normalize-space(translate(text =" f"'{form_label following self::input or self::textarea "

"or self::select][1]"

)

box = page.locator(xpath).first

if box.count == 0:

try: # fallback: proper <label for

box = page.get_by_label(form_label, exact=True).first box.wait_for(timeout=1000)

except Exception:

return False

tag = box.evaluate("e => e.tagName.toLowerCase if tag == "select":

box.select_option(label=value)

else:

box.fill(value)

return True

def main :

pdfs = sorted(PDF_FOLDER.glob pdf"))

if not pdfs:

sys.exit(f"No PDFs found in '{PDF_FOLDER )

with sync_playwright as p:

browser = p.chromium.launch(headless=False)

page = browser.new_context .new_page page.goto(SITE_URL)

input("Log in in the browser window, then press Enter here... ") for pdf in pdfs:

data = parse_pdf(pdf)

input(f"\nOpen the BLANK form for {pdf.name}, then press Enter... ") missing = []

for label, value in data.items :

if SKIP_NA and value.strip .upper == "N/A":

continue

try:

ok = fill_field(page, label, value)

except Exception:

ok = False

if not ok:

missing.append(label)

print(f"{pdf.name}: filled {len(data) - len(missing)}/{len(data)}") if missing:

print(" Could not fill (do by hand):", ", ".join(missing)) if SUBMIT_AUTOMATICALLY and not missing:

page.get_by_text(SUBMIT_BUTTON_TEXT, exact=True).first.click else:

input("Check the form, submit it yourself, then press Enter... ") browser.close if __name__ == "__main__":

main



Contact this candidate