I'm Lennox Kuria, an elite AI Evaluation Engineer & Algorithmic Reasoning Specialist with proven expertise in designing adversarial test cases, constructing deterministic evaluation rubrics, and performing surgical code audits for frontier LLM training pipelines. I have a deep command of Advanced Data Structures & Algorithms (Tries, DAGs, Dynamic Programming), high-concurrency Python architectures, and formal verification methodologies (Lean, TLA+). I specialize in breaking AI reasoning pathways through systematic edge-case construction, repo-wide code evaluation, and multi-step agentic workflow validation. I've evaluated 1000+ AI outputs across software engineering, finance, mathematics, and Web3 domains, maintaining 94% peer reviewer agreement on complex RLHF tasks. I design robust evaluation rubrics for code correctness, security vulnerabilities, and maintainability across multi-file refactors and architectural migrations. Authorized to work as a B2B/C2C contractor for US entities.

Lennox Kuria

I'm Lennox Kuria, an elite AI Evaluation Engineer & Algorithmic Reasoning Specialist with proven expertise in designing adversarial test cases, constructing deterministic evaluation rubrics, and performing surgical code audits for frontier LLM training pipelines. I have a deep command of Advanced Data Structures & Algorithms (Tries, DAGs, Dynamic Programming), high-concurrency Python architectures, and formal verification methodologies (Lean, TLA+). I specialize in breaking AI reasoning pathways through systematic edge-case construction, repo-wide code evaluation, and multi-step agentic workflow validation. I've evaluated 1000+ AI outputs across software engineering, finance, mathematics, and Web3 domains, maintaining 94% peer reviewer agreement on complex RLHF tasks. I design robust evaluation rubrics for code correctness, security vulnerabilities, and maintainability across multi-file refactors and architectural migrations. Authorized to work as a B2B/C2C contractor for US entities.

Available to hire

I’m Lennox Kuria, an elite AI Evaluation Engineer & Algorithmic Reasoning Specialist with proven expertise in designing adversarial test cases, constructing deterministic evaluation rubrics, and performing surgical code audits for frontier LLM training pipelines. I have a deep command of Advanced Data Structures & Algorithms (Tries, DAGs, Dynamic Programming), high-concurrency Python architectures, and formal verification methodologies (Lean, TLA+). I specialize in breaking AI reasoning pathways through systematic edge-case construction, repo-wide code evaluation, and multi-step agentic workflow validation.

I’ve evaluated 1000+ AI outputs across software engineering, finance, mathematics, and Web3 domains, maintaining 94% peer reviewer agreement on complex RLHF tasks. I design robust evaluation rubrics for code correctness, security vulnerabilities, and maintainability across multi-file refactors and architectural migrations. Authorized to work as a B2B/C2C contractor for US entities.

See more

Experience Level

Expert
Expert
Expert
Expert
Expert
Expert
Expert
Expert
Expert
Expert
Expert
Expert
Expert
Expert
Expert
Expert
Expert
Expert
Expert
Expert
Expert
Expert
Expert
See more

Language

English
Fluent

Work Experience

Senior AI Training & Evaluation Specialist at Independent AI Contractor / Datacurve AI Shipd|Contract
December 1, 2025 - Present
Architect deterministic evaluation environments and generate high-fidelity RLHF datasets for frontier LLMs; evaluate 200+ AI outputs monthly; maintain 94% agreement with peer reviewers on complex RLHF tasks; design adversarial 'Quests' to expose weaknesses in repo-wide code generation and multi-agent workflows; perform rigorous code QA and PR reviews for large-scale distributed systems, identifying subtle concurrency bugs and optimizing algorithms while enforcing strict typing.
Senior AI Training & Evaluation Specialist at Datacurve AI
December 1, 2025 - Present
Architect deterministic evaluation environments and generate high-fidelity RLHF datasets for frontier LLMs; evaluate 200+ AI outputs monthly across software engineering, finance, mathematics, and Web3; construct adversarial quests to expose weaknesses in repo-wide code generation and agentic workflows; conduct rigorous code QA and PR reviews for large-scale distributed systems; design evaluation rubrics for code correctness, security vulnerabilities, and maintainability across multi-file refactors and architectural migrations.
AI Systems Evaluator / RLHF Engineer at Independent / Stealth Labs|Contract
February 1, 2025 - August 1, 2025
Build deterministic Docker environments and Python AST verifiers to evaluate frontier AI models for production reliability; generate specialized training datasets focused on edge computing, adversarial prompt defense, and multi-agent coordination; validate mathematical proofs using Lean, assessing 200+ AI-generated proofs; identify gaps between syntactic correctness and problem-solving.
AI Systems Evaluator / RLHF Engineer at Stealth Labs
February 1, 2025 - August 1, 2025
Build deterministic Docker environments and Python-based AST verifiers to evaluate frontier AI models for production reliability; generate specialized training datasets focused on edge computing scenarios, adversarial prompt defense, and multi-agent system coordination; validate mathematical proofs using Lean framework, assessing 200+ AI-generated proofs across number theory, graph theory, and combinatorics; identify gaps between syntactically correct proofs and engineering problems.
Senior Full Stack Software Engineer at Moovx
November 1, 2023 - January 1, 2025
Designed highly scalable web applications bridging modern reactive frontends (React, Vue.js) with secure, object-oriented backend APIs handling 100K+ daily requests; led architectural design for cloud migrations using strangler-fig patterns to safely transition legacy monoliths into microservices; built and maintained robust CI/CD pipelines with Terraform and Kubernetes, enforcing blue-green deployments; architected scalable stream processing and data orchestration pipelines (Apache Airflow, Golang) capable of securely processing 10M+ payloads daily.
Lead Architect & Developer at Aegis-Market (Independent Project)|Apprenticeship
February 1, 2023 - March 1, 2026
Architected autonomous multi-agent arbitrage and negotiation engine; designed deterministic data validation pipelines using JSON schemas; developed RAG (Retrieval-Augmented Generation) pipeline to ground agent reasoning in verified internal data; enforced OpSec with isolated containerized execution environments preventing unverified agents from autonomous state changes without human approval.
Lead Architect & Developer at Aegis-Market
February 1, 2023 - March 1, 2026
Architected autonomous multi-agent arbitrage and negotiation engine using FastAPI for high-throughput backend orchestration; designed strict deterministic data validation pipelines using Pydantic, ensuring all agent inputs/outputs adhere to rigid JSON schemas before execution; developed Retrieval-Augmented Generation (RAG) pipeline to ground agent reasoning in verified internal data with strict cost-control mechanisms; enforced operational security by designing isolated, containerized execution environments preventing unverified agents from making autonomous state changes without human approval.
Data Infrastructure & Systems Analyst at Freelance Tech Solutions
March 1, 2017 - December 1, 2020
Built and maintained complex distributed systems and backend workflows using Python, ensuring structural integrity for large-scale enterprise data integrations; developed full-stack features from database schema design to client-side UI; integrated automated testing frameworks and version control to reduce post-release hotfixes; created backend verifiers that autonomously monitored data quality and pipeline health.

Education

Bachelor of Science in Computer Science at University of California, Berkeley
January 11, 2030 - January 1, 2023
Bachelor of Science in Computer Science at University of California, Berkeley
January 11, 2030 - January 1, 2023
Bachelor of Science in Computer Science at University of California, Berkeley
January 11, 2030 - January 1, 2023

Qualifications

Add your qualifications or awards here.

Industry Experience

Software & Internet, Financial Services, Professional Services, Media & Entertainment, Healthcare