Post Job Free
Sign in

AI Trainer & Red-Team Evaluator

Location:
Sacramento, CA
Salary:
$40/hr
Posted:
June 30, 2026

Contact this candidate

Resume:

NIKITA KOVALENKO

SKILLS AI & Model Evaluation: red-teaming, trust & safety,

adversarial testing, QA evaluation

Data & Analysis: structured annotation

(JSON/taxonomies), signal vs noise detection,

pattern recognition

Content & Linguistics: narrative analysis, multilingual evaluation, Russian/English fluency

Reporting: insight synthesis, error analysis,

executive-level reporting

Platforms: social media ecosystems, UGC

datasets, cross-platform AI systems

Languages: English (Native), Russian (Native)

EDUCATION August 2017 – May 2019

Studied at Sacramento State University for two years; program unfinished due to relocation overseas. Sacramento State University, California — Mechanical Engineering SUMMARY AI Trainer and Data Annotator with experience in model evaluation, safety testing, and NLP data workflows. Strong background in QA, error analysis, and identifying model failure patterns to improve performance and reliability. Experienced in fast-paced, experimental AI environments including red- teaming and multimodal dataset work. Bilingual in English and Russian with cross-cultural communication experience.

Address: Sacramento, CA 95831 Phone: 279-***-**** Email: ****************@*****.*** WORK

EXPERIENCE

Stealth AI Company — Data Annotator September 2025 – March 2026 Parsed long-form meeting transcripts to extract structured context items, classifying tasks, decisions, and discussions using consistent schema logic.

Verified timestamps, speakers, and responsibilities across extracted items to ensure data integrity for downstream automation and analytics.

Collaborated with QA teams to improve annotation guidelines, proposing new edge-case categories and providing pattern-based insights that directly shaped how the annotation framework evolved. Mercor — Audio Model Trainer August 2025 – November 2025 Recorded voice samples for multimodal AI model training, consistently meeting technical standards for clarity, pacing, pronunciation, and audio quality. Produced spoken descriptions of visual content following strict linguistic and stylistic guidelines to support dataset consistency.

Reviewed and evaluated audio clips for quality control, flagging issues that impacted model performance. Outlier.ai — Russian AI Trainer July 2024 – November 2025 Led Russian-language model evaluation, fact-checking responses for accuracy and flagging reliability issues across a high volume of outputs.

Rated and ranked AI responses with a critical eye, going beyond surface-level scoring to identify where evaluation parameters were being gamed or missed.

Audited AI-generated voice responses for quality, catching issues with naturalness, pacing, and adherence to requirements that automated checks overlooked. Brought deep knowledge of Russian language nuance, cultural context, and idiomatic expression to improve model outputs in ways native speakers would actually notice. Designed task-specific prompts that pushed model performance, iterating based on observed behavior rather than just following templates.

Tracked behavioral patterns across model outputs and translated findings into concrete improvement recommendations for the broader team.

Invisible Technologies, Inc. — Advanced AI Data Trainer

(Red Teaming), Safety AI Data Trainer

May 2023 – September 2025

Participated in model development workflows spanning mid-training, post-training, and evaluation stages. Took ownership of prompt design in high-risk scenarios, actively probing model behavior around minor safety, CSAM/CSA, PCG, and violent content to surface weaknesses before they reached production. Embedded with the Trust & Safety team to audit model responses against safety policies, flagging gaps and pushing for stronger content moderation criteria. Sought out sensitive, unsafe, and adversarial edge cases to stress-test model limits and feed red-teaming strategies with concrete, pattern-based findings.

Drove PSC and SC evaluation cycles by surfacing hard edge cases and translating QA insights directly into fine-tuning recommendations.



Contact this candidate