AI engineer who authors and verifies high-difficulty agentic benchmark tasks for frontier-model post-training and evaluation, owning environment design, programmatic verifiers, containerization, reward-hack hardening, and pass@k measurement end to end. Establishes task difficulty empirically against frontier reasoning models while also shipping production ML systems. MCA candidate (June 2026) with experience building reproducible benchmark environments and robust evaluation harnesses, plus production-ready ML platforms and autonomous multi-agent pipelines with streaming logs and provider fallback.

Chandrasekhar Thapa

AI engineer who authors and verifies high-difficulty agentic benchmark tasks for frontier-model post-training and evaluation, owning environment design, programmatic verifiers, containerization, reward-hack hardening, and pass@k measurement end to end. Establishes task difficulty empirically against frontier reasoning models while also shipping production ML systems. MCA candidate (June 2026) with experience building reproducible benchmark environments and robust evaluation harnesses, plus production-ready ML platforms and autonomous multi-agent pipelines with streaming logs and provider fallback.

Available to hire

AI engineer who authors and verifies high-difficulty agentic benchmark tasks for frontier-model post-training and evaluation, owning environment design, programmatic verifiers, containerization, reward-hack hardening, and pass@k measurement end to end. Establishes task difficulty empirically against frontier reasoning models while also shipping production ML systems.

MCA candidate (June 2026) with experience building reproducible benchmark environments and robust evaluation harnesses, plus production-ready ML platforms and autonomous multi-agent pipelines with streaming logs and provider fallback.

See more

Experience Level

Expert
Expert
Expert
Intermediate
Intermediate
Intermediate
Intermediate
See more

Language

Work Experience

AI Engineer (Contract) — Agentic Benchmark & Evaluation Authoring at Handshake AI — Dynamo Program 2026
January 1, 2026 - Present
Author agentic-coding benchmark tasks for frontier-model post-training and evaluation. Each task is shipped as a reproducible Docker environment with an all-or-nothing reward contract and a programmatic verifier. Work spans debugging-and-repair, query optimization, and RL-infrastructure domains. Built containerized verification harnesses (including build integrity and containment), implemented reward-hacking batteries, and hardened solution-leakage vectors at the image and test-execution layers. Established empirical difficulty using measurement (trajectory/failure analysis, per-trap defeat probability modeling for pass@k forecasting, mutation testing to ensure requirements are load-bearing) and ran controlled disclosure A/B experiments to validate solve-rate improvements.

Education

Master of Computer Applications (MCA) at Kalinga Institute of Industrial Technology (KIIT), Bhubaneswar
January 11, 2030 - June 1, 2026
Bachelor of Science in Computer Science at Ravenshaw University, Cuttack
January 1, 2020 - July 1, 2023

Qualifications

Ethical Hacking with AI (Internshala) — OWASP Top 10, VAPT automation and AI-assisted security testing
January 11, 2030 - August 27, 2026
Alteryx Foundational Micro-Credential
January 11, 2030 - August 27, 2026

Industry Experience

Software & Internet, Computers & Electronics, Education, Professional Services