AI Model Evaluator with 6+ years of remote experience supporting LLM training, RLHF, prompt evaluation, response ranking, data annotation, and software engineering review. I assess AI outputs for reasoning quality, factuality, safety, instruction following, and code correctness, with strong attention to detail and consistent high-quality results in asynchronous, guideline-driven environments. I combine a background in Computer Science, Cybersecurity, and Software Engineering with strong Python/SQL debugging and analytical skills to produce structured feedback, reference answers, and QA checks that help improve model performance and reliability.

ALEX WACHIRA

AI Model Evaluator with 6+ years of remote experience supporting LLM training, RLHF, prompt evaluation, response ranking, data annotation, and software engineering review. I assess AI outputs for reasoning quality, factuality, safety, instruction following, and code correctness, with strong attention to detail and consistent high-quality results in asynchronous, guideline-driven environments. I combine a background in Computer Science, Cybersecurity, and Software Engineering with strong Python/SQL debugging and analytical skills to produce structured feedback, reference answers, and QA checks that help improve model performance and reliability.

Available to hire

AI Model Evaluator with 6+ years of remote experience supporting LLM training, RLHF, prompt evaluation, response ranking, data annotation, and software engineering review. I assess AI outputs for reasoning quality, factuality, safety, instruction following, and code correctness, with strong attention to detail and consistent high-quality results in asynchronous, guideline-driven environments.

I combine a background in Computer Science, Cybersecurity, and Software Engineering with strong Python/SQL debugging and analytical skills to produce structured feedback, reference answers, and QA checks that help improve model performance and reliability.

See more

Language

Work Experience

Technical AI Researcher & Content Writer
January 1, 2020 - Present
Research emerging AI technologies, LLM capabilities, and evaluation methodologies. Produce technical documentation and educational AI content to support learning and adoption of best practices in AI evaluation.
Software Engineering Reviewer
January 1, 2020 - Present
Review Python and SQL code for correctness, maintainability, edge cases, efficiency, and security. Evaluate AI-generated code explanations and debugging workflows. Assess algorithms, data structures, APIs, and software design decisions to ensure technical soundness and secure, reliable implementation.
Data Annotation Specialist
January 1, 2019 - Present
Annotate text, code, and structured datasets for machine learning projects. Perform quality audits and ensure compliance with annotation guidelines. Support taxonomy development and maintain dataset consistency across projects.
Remote AI Trainer & AI Model Evaluator
January 1, 2019 - Present
Evaluate AI-generated responses across coding, mathematics, reasoning, cybersecurity, writing, and general knowledge. Rank multiple model outputs using detailed rubrics and RLHF-style evaluation criteria. Detect hallucinations, logical inconsistencies, factual inaccuracies, unsafe responses, and formatting issues. Create reference answers and structured feedback used to improve large language model performance. Work independently in remote, guideline-driven environments while maintaining strong attention to detail.

Education

Add your educational history here.

Qualifications

Master's Degree in Computer Science - Cybersecurity
January 1, 2017 - January 1, 2018
Bachelor's Degree in Computer Science - Software Engineering
January 1, 2013 - January 1, 2016

Industry Experience

Software & Internet, Professional Services, Education, Computers & Electronics