I’m Ndegwa Mariga Stephen, an AI Engineer and Data Scientist based in Nairobi, Kenya. I combine strong mathematical and statistical foundations with hands-on engineering to build reliable AI systems across machine learning, deep learning, NLP, RAG, and LLM evaluation. I enjoy turning complex evaluation needs into reproducible workflows and practical software that teams can depend on. In my recent work, I’ve designed benchmark tasks for evaluating LLMs, built validation workflows to ensure correctness and reproducibility, and performed high-quality evaluation of AI outputs (including robotics video annotation with 96% accuracy). I also develop AI-powered applications using FastAPI, PostgreSQL/pgvector, and Docker—delivering context-aware tutoring and research assistant features that balance performance and cost while staying focused on quality.

Ndegwa Mariga Stephen

I’m Ndegwa Mariga Stephen, an AI Engineer and Data Scientist based in Nairobi, Kenya. I combine strong mathematical and statistical foundations with hands-on engineering to build reliable AI systems across machine learning, deep learning, NLP, RAG, and LLM evaluation. I enjoy turning complex evaluation needs into reproducible workflows and practical software that teams can depend on. In my recent work, I’ve designed benchmark tasks for evaluating LLMs, built validation workflows to ensure correctness and reproducibility, and performed high-quality evaluation of AI outputs (including robotics video annotation with 96% accuracy). I also develop AI-powered applications using FastAPI, PostgreSQL/pgvector, and Docker—delivering context-aware tutoring and research assistant features that balance performance and cost while staying focused on quality.

Available to hire

I’m Ndegwa Mariga Stephen, an AI Engineer and Data Scientist based in Nairobi, Kenya. I combine strong mathematical and statistical foundations with hands-on engineering to build reliable AI systems across machine learning, deep learning, NLP, RAG, and LLM evaluation. I enjoy turning complex evaluation needs into reproducible workflows and practical software that teams can depend on.

In my recent work, I’ve designed benchmark tasks for evaluating LLMs, built validation workflows to ensure correctness and reproducibility, and performed high-quality evaluation of AI outputs (including robotics video annotation with 96% accuracy). I also develop AI-powered applications using FastAPI, PostgreSQL/pgvector, and Docker—delivering context-aware tutoring and research assistant features that balance performance and cost while staying focused on quality.

See more

Work Experience

AI Repository Development Specialist at AfterQuery
May 1, 2026 - June 1, 2026
Developed production-style software repositories and realistic Git histories for AI evaluation environments. Configured reproducible development environments, testing workflows, and technical documentation. Debugged software defects and validated repositories against automated evaluation requirements, producing clean and maintainable code following software engineering best practices.
AI Training Contributor / AI Benchmarking Specialist at AfterQuery
May 1, 2026 - July 1, 2026
Designed benchmark tasks for training and evaluating LLMs across software engineering, DevOps, and AI domains. Authored technical specifications, evaluation criteria, and reproducible coding environments to support consistent AI model evaluation. Built validation workflows and executed local tests to verify benchmark correctness, reproducibility, and expected behavior; identified edge cases, dependency issues, and debugging scenarios to improve benchmark quality and reliability.
Data Annotation Consultant at IMerit Scholars
March 1, 2026 - April 1, 2026
Evaluated AI-generated outputs against detailed quality guidelines and used analytical judgment to improve annotation quality. Contributed to AI model improvement through systematic evaluation and quality assurance in a fast-paced environment.
Data Scientist / Data Analyst
January 1, 2022 - January 1, 2025
Built end-to-end data analytics pipelines using Python, R, and SQL for data preparation, exploratory analysis, and predictive modeling. Developed predictive models and statistical analyses to support decision-making. Created dashboards and visualizations for technical and non-technical stakeholders, automated data preparation, and maintained reproducible analytical workflows.

Education

Bachelor of Science in Mathematics (Statistics) at University of Nairobi
January 11, 2030 - September 1, 2025

Qualifications

Advanced Python Programming
January 11, 2030 - September 1, 2026
SQL for Data Science
January 11, 2030 - September 1, 2026
Machine Learning
January 11, 2030 - September 1, 2026
Data Visualization with Tableau
January 11, 2030 - September 1, 2026
Prompt Engineering & AI Tools
January 11, 2030 - September 1, 2026

Industry Experience

Education, Software & Internet, Computers & Electronics
    Customer Churn Prediction - Machine Learning & Predictive Analytics

    Developed a machine learning pipeline to predict customer churn and identify factors associated with customer attrition. I performed data cleaning, exploratory analysis, feature preparation, model training, and evaluation across multiple classification algorithms. I compared Decision Trees, Random Forest, Gradient Boosting, XGBoost, and LightGBM models and evaluated their performance to identify a suitable predictive approach. The project strengthened my experience in model comparison, validation, feature engineering, and translating predictive modeling into actionable business insights.

    Project Pluto - AI Benchmark Development & LLM Evaluation

    Contributed to AfterQuery’s Project Pluto, developing high-quality benchmark tasks designed for training and evaluating large language models. I created realistic software engineering, DevOps, data engineering, and AI scenarios, wrote detailed task specifications and evaluation criteria, and developed reproducible environments and validation workflows. I executed tests locally, investigated edge cases and dependency issues, and ensured submissions were technically correct, reproducible, and compatible with automated evaluation pipelines. The work required structured reasoning, attention to detail, technical documentation, software debugging, and independent execution in a remote asynchronous environment.

    Project Silver - AI Training Repository Development

    Contributed to AfterQuery’s Project Silver by developing realistic software repositories designed for AI training and evaluation. I built production-style codebases with authentic Git histories, implemented features according to benchmark specifications, configured dependencies and testing workflows, and validated repositories against automated evaluation systems. I also diagnosed software defects, dependency problems, and configuration issues while maintaining clean, reproducible, and well-documented repositories. This project strengthened my ability to combine software engineering judgment with the requirements of AI training and evaluation.

    LLM Playground - LLM Application & API Integration

    Developed an LLM experimentation platform for interacting with language models through a web interface and backend API. I designed the application architecture using Next.js and FastAPI, integrated local LLM inference, and implemented API-based communication between the frontend and backend. The project provided practical experience in LLM application development, API design, asynchronous workflows, model integration, and building user-facing AI interfaces.

    Marixion - AI & Digital Solutions Platform

    Built and contributed to digital and AI-powered solutions through Marixion, combining software engineering, AI, data, and modern web technologies to address practical business problems. My work has involved developing web applications, backend services, APIs, database integrations, and AI-enabled functionality. I have worked across the development lifecycle, including requirements analysis, architecture, implementation, debugging, testing, deployment, and technical documentation. Available online at marixion.com

    MedAITUTOR — AI-Powered Learning & Document Intelligence Platform

    MEDAITUTOR is an AI-powered learning platform designed to help students interact intelligently with their study materials. I developed the backend using Python and FastAPI and implemented a RAG pipeline that processes uploaded documents, generates embeddings, performs semantic retrieval, and provides context-aware responses. I also developed AI-powered features such as document-based chat and automated flashcard generation, integrating the application with PostgreSQL/Supabase and vector search. The project involved designing APIs, debugging backend services, integrating LLM capabilities, and building a practical AI application from concept to working system. It’s publicly available at medaitutor.com