I’m Helen Badmus, an AI evaluation and quality specialist with hands-on experience assessing LLM outputs and training data for safety, factual accuracy, logical reasoning, and policy compliance. I’ve evaluated 10,000+ AI-generated responses across multi-turn conversations, using structured rubrics and edge-case analysis to help improve model alignment. Alongside my AI work, I build and test user-facing web experiences by translating AI and evaluation requirements into responsive HTML/CSS/JavaScript interfaces. I also specialize in speech and audio dataset validation, and I contribute to conversational data creation and RLHF-style evaluation workflows that strengthen real-world AI behavior.

Helen Badmus

I’m Helen Badmus, an AI evaluation and quality specialist with hands-on experience assessing LLM outputs and training data for safety, factual accuracy, logical reasoning, and policy compliance. I’ve evaluated 10,000+ AI-generated responses across multi-turn conversations, using structured rubrics and edge-case analysis to help improve model alignment. Alongside my AI work, I build and test user-facing web experiences by translating AI and evaluation requirements into responsive HTML/CSS/JavaScript interfaces. I also specialize in speech and audio dataset validation, and I contribute to conversational data creation and RLHF-style evaluation workflows that strengthen real-world AI behavior.

Available to hire

I’m Helen Badmus, an AI evaluation and quality specialist with hands-on experience assessing LLM outputs and training data for safety, factual accuracy, logical reasoning, and policy compliance. I’ve evaluated 10,000+ AI-generated responses across multi-turn conversations, using structured rubrics and edge-case analysis to help improve model alignment.

Alongside my AI work, I build and test user-facing web experiences by translating AI and evaluation requirements into responsive HTML/CSS/JavaScript interfaces. I also specialize in speech and audio dataset validation, and I contribute to conversational data creation and RLHF-style evaluation workflows that strengthen real-world AI behavior.

See more

Language

English
Fluent
Yoruba
Fluent
French
Beginner

Work Experience

Front-End AI Engineering Intern at FlyRank AI
July 1, 2026 - Present
Translate AI and evaluation requirements into functional, responsive web interfaces using HTML, CSS, and JavaScript. Develop, test, and debug front-end code supporting integration of AI models into user-facing applications. Perform functional, responsiveness, and usability testing, and use modern AI-assisted development workflows to improve UI quality and troubleshooting.
AI Data Annotator & Content Reviewer at Atlas Capture AI
August 1, 2025 - June 1, 2026
Evaluate multi-turn conversations and AI-generated responses for factual correctness, logical coherence, relevance, and natural language quality. Perform multi-turn annotation, ranking, and RLHF tasks aligned to project-specific guidelines. Identify response-quality issues and edge cases requiring additional review or correction across large volumes of content.
AI Tutor & Data Annotator (Contract) at Mindrift
January 1, 2025 - Present
Evaluate large language model responses for safety, factual accuracy, relevance, and adherence to human evaluation guidelines. Identify ethical, cultural, and safety risks and provide structured assessments to support model alignment. Create and evaluate multi-turn conversational data for AI agents/storefront contexts, analyzing instruction following, context awareness, naturalness, and edge cases.
AI Data annotation specialist at AtlasCapture AI
January 1, 2025 - Present
Labeled and validated text and document datasets for AI projects. Ensured accuracy, completeness, and compliance with data privacy standards. Conducted quality checks to maintain high data quality for AI training.
Virtual Assistant & Content Coordinator at Remote
January 1, 2024 - Present
Coordinate digital assets, content calendars, and administrative workflows using spreadsheets and productivity tools. Organize recurring workflows to improve operational efficiency and maintain accurate records, supporting content and administrative tasks across remote work environments.
AI Data Contributor at Quikrsignal.ai
January 1, 2024 - Present
Capture and submit targeted video and audio datasets for computer vision and audio machine learning projects. Follow collection and quality-control guidelines to ensure submissions meet project requirements. Review captured data for accuracy, completeness, and compliance before submission.
Acoustic Data Specialist (Project-Based) at Uber AI
April 1, 2023 - Present
Record, review, and validate speech and audio datasets used for speech recognition and text-to-speech model development. Apply strict data-quality requirements during collection and audio sample validation. Document data inconsistencies, edge cases, and safety compliance issues to support continuous dataset and model improvement.
Quality Assurance Tester (Freelance) at uTest & Test IO
January 1, 2023 - March 1, 2026
Conduct exploratory, functional, and regression testing across mobile and web applications. Identify defects and document clear reproduction steps, supported with screenshots and system logs. Evaluate behavior across scenarios and user flows while following client-specific testing requirements.
AI Data Annotation & QA specialist at Scale AI
March 1, 2022 - March 1, 2025
Annotated and evaluated datasets for AI training and model improvement. Reviewed AI outputs for accuracy, safety, and guideline compliance. Conducted quality assurance checks and flagged errors or inconsistencies.
Freelance Content Reviewer and Researcher at Upwork
January 1, 2019 - Present
Reviewed and edited content with attention to clarity, accuracy, and compliance. Achieved 95% job success rate with consistent five-star client reviews.

Education

Bachelor's degree at KWARA STATE UNIVERSITY
January 1, 2019 - January 1, 2023
Master's degree at KWARA STATE UNIVERSITY
January 1, 2023 - April 12, 2026
B.Sc. Computer and Data Science at Kwara State University
January 11, 2030 - November 1, 2023

Qualifications

Microsoft Azure AI Essentials: Workloads and Machine Learning on Azure
August 1, 2026 - September 1, 2026
AI Fluency: Framework and Foundations
January 1, 2026 - September 1, 2026
Social Media Management Blueprint Certificate
January 1, 2022 - September 1, 2026

Industry Experience

Software & Internet, Professional Services, Other, Computers & Electronics, Education, Media & Entertainment
    High-Quality Data Annotation for Machine Learning Training

    This project involved preparing and labeling structured datasets for machine learning model training. The goal was to ensure high-quality, consistent, and accurately annotated data that can be used to improve the performance of AI systems.

    I worked with both image and text-based datasets, applying clear labeling guidelines to ensure consistency, relevance, and precision across all entries. Each data point was carefully reviewed to reduce noise, eliminate ambiguity, and maintain dataset integrity.

    Key responsibilities included:

    Annotating and labeling image and text data according to defined guidelines

    Ensuring consistency and accuracy across all labeled datasets

    Identifying and correcting ambiguous or incorrectly tagged data

    Maintaining high standards of data quality for machine learning use cases

    Structuring datasets for improved model training efficiency

    Outcome:
    The final annotated datasets were clean, consistent, and optimized for machine learning model training, improving their usability for supervised learning tasks and reducing potential errors in model outputs.

    AI Response Evaluation and Quality Assurance

    Project overview
    This project involved evaluating AI-generated responses for accuracy, clarity, coherence, and overall quality. The goal was to identify weaknesses in model outputs and improve them to meet human-level communication standards.

    Approach
    I reviewed multiple AI responses across different prompts, assessing them for factual correctness, logical consistency, tone appropriateness, and grammatical accuracy. Where necessary, I rewrote responses to improve structure, reasoning, and readability.

    Key Responsibilities
    Detected factual errors and logical inconsistencies in AI outputs
    Improved clarity, coherence, and readability of responses
    Ensured tone alignment with user intent and context
    Rewrote and refined outputs for better logical flow and precision

    Outcome
    The refined responses showed improved accuracy, stronger reasoning, and more natural communication, making them better aligned with professional AI training and evaluation standards.