Core Contributor at Fleet
February 2, 2026 - June 1, 2026Selected as one of approximately 10 Core Contributors from a workforce of 800+ evaluators, based on output quality, reliability,
and depth of analysis
• Evaluated AI agents completing multi-step enterprise workflows at 40+ hours/week, working inside simulated environments
replicating Salesforce, Workday, QuickBooks, Outlook, BI dashboards, and DBT pipelines
• Designed and refined adversarial evaluation prompts across complex multi-step workflows, requiring agents to discover
information, follow sequential instructions, make decisions, and correctly apply earlier findings; surfaced failure types spanning
hallucination, verifier misalignment, instruction drift, and incomplete action chains
• Reviewed model outputs and workflow traces for hallucinations, logical inconsistencies, verifier alignment, and recurring failure
patterns
• Worked directly with JSON outputs and SQL-based interfaces to assess structured data quality, reasoning accuracy, and instruction
adherence
• Identified broken navigation, missing actions, weak validation, unrealistic seeded data, and UI gaps that would block successful
task completion
• Delivered structured QA reports translating technical findings into clear, actionable feedback for model improvement teams
• Selected for OpenClaw, Fleet's advanced agentic evaluation project, responsible for designing and building seed worlds: the
simulated enterprise environments used to stress-test AI agent behavior across complex, multi-step workflows spanning Salesforce,
Workday, QuickBooks, Outlook, BI dashboards, and DBT pipelines
• Served on a small pre-launch environment QA team to review simulated enterprise environments before release to the broader
annotator workforce, surfacing data integrity issues, broken workflows, unrealistic content, and environment-level bugs that would
have compromised evaluation quality downstream