I support AI training pipelines by applying project-specific guidelines across text, audio, and multimodal datasets—helping teams produce accurate, well-structured training data. My work includes careful evaluation and red-teaming, identifying edge cases, and analyzing model failure patterns to improve instruction-following and reduce hallucinations. I also focus on high-quality annotation and QA, including millisecond timestamp alignment for audio/video, speech transcription, and labeling emotional and vocal attributes. I collaborate across enterprise tools and workflows (Jira, Slack, Google Workspace, Excel, and SQL-based interfaces) and provide structured, stakeholder-ready reporting that documents both quality findings and user impact.…

Nancy Power

I support AI training pipelines by applying project-specific guidelines across text, audio, and multimodal datasets—helping teams produce accurate, well-structured training data. My work includes careful evaluation and red-teaming, identifying edge cases, and analyzing model failure patterns to improve instruction-following and reduce hallucinations. I also focus on high-quality annotation and QA, including millisecond timestamp alignment for audio/video, speech transcription, and labeling emotional and vocal attributes. I collaborate across enterprise tools and workflows (Jira, Slack, Google Workspace, Excel, and SQL-based interfaces) and provide structured, stakeholder-ready reporting that documents both quality findings and user impact.…

Available to hire

I support AI training pipelines by applying project-specific guidelines across text, audio, and multimodal datasets—helping teams produce accurate, well-structured training data. My work includes careful evaluation and red-teaming, identifying edge cases, and analyzing model failure patterns to improve instruction-following and reduce hallucinations.

I also focus on high-quality annotation and QA, including millisecond timestamp alignment for audio/video, speech transcription, and labeling emotional and vocal attributes. I collaborate across enterprise tools and workflows (Jira, Slack, Google Workspace, Excel, and SQL-based interfaces) and provide structured, stakeholder-ready reporting that documents both quality findings and user impact.

See more

Work Experience

Core Contributor at Fleet
February 2, 2026 - June 1, 2026
Selected as one of approximately 10 Core Contributors from a workforce of 800+ evaluators, based on output quality, reliability, and depth of analysis • Evaluated AI agents completing multi-step enterprise workflows at 40+ hours/week, working inside simulated environments replicating Salesforce, Workday, QuickBooks, Outlook, BI dashboards, and DBT pipelines • Designed and refined adversarial evaluation prompts across complex multi-step workflows, requiring agents to discover information, follow sequential instructions, make decisions, and correctly apply earlier findings; surfaced failure types spanning hallucination, verifier misalignment, instruction drift, and incomplete action chains • Reviewed model outputs and workflow traces for hallucinations, logical inconsistencies, verifier alignment, and recurring failure patterns • Worked directly with JSON outputs and SQL-based interfaces to assess structured data quality, reasoning accuracy, and instruction adherence • Identified broken navigation, missing actions, weak validation, unrealistic seeded data, and UI gaps that would block successful task completion • Delivered structured QA reports translating technical findings into clear, actionable feedback for model improvement teams • Selected for OpenClaw, Fleet's advanced agentic evaluation project, responsible for designing and building seed worlds: the simulated enterprise environments used to stress-test AI agent behavior across complex, multi-step workflows spanning Salesforce, Workday, QuickBooks, Outlook, BI dashboards, and DBT pipelines • Served on a small pre-launch environment QA team to review simulated enterprise environments before release to the broader annotator workforce, surfacing data integrity issues, broken workflows, unrealistic content, and environment-level bugs that would have compromised evaluation quality downstream
AI Data Evaluator at Telus
January 1, 2026 - Present
Evaluated AI-generated content and model outputs across text and multimodal tasks for quality, accuracy, and safety compliance • Applied rubric-based frameworks to assess instruction-following, coherence, and response usefulness • Flagged edge cases and systemic failure patterns to support ongoing model improvement
Trainor/Quality Auditor at Mercor
August 2, 2025 - Present
Evaluated AI-generated outputs across text, image, audio, and video modalities using comparative ranking and rubric-based auditing • Conducted millisecond-level audio transcription alignment and labelled emotional tone, vocal attributes, and speech quality for AI training datasets • Identified labelling inconsistencies and systemic model errors to improve training data reliability

Education

Diploma at Durham College
January 11, 2030 - November 10, 2025
Diploma at Seneca College
January 11, 2030 - November 10, 2025

Qualifications

QuickBooks Certified
January 1, 2024 - November 10, 2025
Bookkeeping Fundamentals
January 1, 2024 - November 10, 2025
QuickBooks Certified
January 1, 2024 - September 17, 2026
Bookkeeping Fundamentals
January 1, 2024 - September 17, 2026
QuickBooks Certified
January 1, 2024 - September 17, 2026
Bookkeeping Fundamentals
January 1, 2024 - September 17, 2026

Industry Experience

Professional Services, Retail, Software & Internet, Media & Entertainment, Other, Financial Services, Telecommunications