I’m an AI evaluation specialist with 4+ years of experience improving language and vision models through rigorous, feedback-driven assessment. I focus on evaluating response quality and safety (helpfulness, harmlessness, and honesty), identifying hallucinations, and ensuring outputs comply with policy and platform standards. I also contribute to high-quality multimodal data work—annotating audio (transcription and speaker diarization), labeling image and video datasets, and supporting large-scale training with careful attention to detail. I enjoy structured reasoning, systematic evaluation, and providing actionable recommendations that help teams and models perform better over time.

Sara Hwanjiru

I’m an AI evaluation specialist with 4+ years of experience improving language and vision models through rigorous, feedback-driven assessment. I focus on evaluating response quality and safety (helpfulness, harmlessness, and honesty), identifying hallucinations, and ensuring outputs comply with policy and platform standards. I also contribute to high-quality multimodal data work—annotating audio (transcription and speaker diarization), labeling image and video datasets, and supporting large-scale training with careful attention to detail. I enjoy structured reasoning, systematic evaluation, and providing actionable recommendations that help teams and models perform better over time.

Available to hire

I’m an AI evaluation specialist with 4+ years of experience improving language and vision models through rigorous, feedback-driven assessment. I focus on evaluating response quality and safety (helpfulness, harmlessness, and honesty), identifying hallucinations, and ensuring outputs comply with policy and platform standards.

I also contribute to high-quality multimodal data work—annotating audio (transcription and speaker diarization), labeling image and video datasets, and supporting large-scale training with careful attention to detail. I enjoy structured reasoning, systematic evaluation, and providing actionable recommendations that help teams and models perform better over time.

See more

Language

English
Fluent
Swahili
Fluent

Work Experience

AI Evaluation Specialist (Remote)
March 1, 2025 - March 1, 2026
Contributed to image generation evaluations by reviewing how well generated images follow prompts, focusing on quality and safety. Performed audio transcription and ASR labeling to support large-scale training of speech recognition models. Carried out speaker diarization tasks by identifying and labeling multiple speakers in audio conversations for voice assistant use cases. Provided verbatim and ASR-corrected transcriptions for phone calls with over 98% accuracy, ensuring transcripts are reliable for downstream tasks.
AI Evaluation / Safety & Quality Rater (Remote)
January 1, 2022 - March 1, 2024
Rated AI model responses on helpfulness, harmlessness, and honesty (RLHF) across large-scale LLM training datasets. Caught and documented hallucinations in AI-generated text, pairing each failure with reference-backed corrections. Evaluated safety and policy adherence of model outputs, detecting unsafe, biased, or non-compliant responses according to platform policies. Assessed code generation responses across multiple programming languages for correctness, efficiency, and style. Conducted response ranking and Likert-scale evaluations to inform reinforcement learning and shape model preferences.
Multimodal Data Annotator (Remote)
January 1, 2021 - March 1, 2022
Labeled video datasets for action recognition, segmentation, and labeling of human activities and interactions to support training of action detection models. Segmented long videos with scene and event labels for video surveillance and sports video analytics. Classified short-form text data (e.g., social media posts, reviews, and headlines) for sentiment and category. Annotated Named Entity Recognition (NER) and Part-of-Speech (POS) tasks on multilingual datasets to support training for search engines and conversational AI. Created image classification labels and generated tight bounding boxes to train computer vision models for retail and e-commerce applications.
Search Relevance & Content Moderation Evaluator (Remote) at Toloka
January 1, 2020 - January 1, 2021
Performed search query intent labeling and search relevance rating to improve search engine ranking and understanding of search underlining algorithms. Worked on ad and content moderation projects by judging user-submitted content based on platform rules to establish brand-safe content. Labeled search relevance pairs and produced relevance score comparisons with text rationale. Judged search query intent across e-commerce, informational, and navigational queries to improve overall relevance.

Education

BSc. Computer Science at Kenyatta University
January 1, 2015 - January 1, 2019
Data Analyst Nanodegree at Udacity
January 1, 2022 - January 1, 2023

Qualifications

Data Analyst Nanodegree
January 1, 2022 - September 16, 2026

Industry Experience

Software & Internet, Computers & Electronics, Media & Entertainment