Hi, I’m Gagandeep S Marigoudar, an IT undergraduate and freelance data specialist focusing on data annotation, AI training data preparation, transcription, and QA. I have 1+ year of hands-on experience working with NLP datasets, RLHF workflows, prompt evaluation, and bilingual data review (English, Kannada, Hindi) for global clients such as Innodata, Appen, and OneForma. I enjoy turning messy data into structured, high-quality datasets that power robust ML and Generative AI systems. I’m proficient with Label Studio, Multimango, Excel; comfortable in Python (Pandas, NumPy), SQL, Tableau/Power BI; I thrive in collaborative teams and love improving model performance through meticulous feedback and data governance.

Gagandeep S Marigoudar

Hi, I’m Gagandeep S Marigoudar, an IT undergraduate and freelance data specialist focusing on data annotation, AI training data preparation, transcription, and QA. I have 1+ year of hands-on experience working with NLP datasets, RLHF workflows, prompt evaluation, and bilingual data review (English, Kannada, Hindi) for global clients such as Innodata, Appen, and OneForma. I enjoy turning messy data into structured, high-quality datasets that power robust ML and Generative AI systems. I’m proficient with Label Studio, Multimango, Excel; comfortable in Python (Pandas, NumPy), SQL, Tableau/Power BI; I thrive in collaborative teams and love improving model performance through meticulous feedback and data governance.

Available to hire

Hi, I’m Gagandeep S Marigoudar, an IT undergraduate and freelance data specialist focusing on data annotation, AI training data preparation, transcription, and QA. I have 1+ year of hands-on experience working with NLP datasets, RLHF workflows, prompt evaluation, and bilingual data review (English, Kannada, Hindi) for global clients such as Innodata, Appen, and OneForma. I enjoy turning messy data into structured, high-quality datasets that power robust ML and Generative AI systems.

I’m proficient with Label Studio, Multimango, Excel; comfortable in Python (Pandas, NumPy), SQL, Tableau/Power BI; I thrive in collaborative teams and love improving model performance through meticulous feedback and data governance.

See more

Language

English
Fluent
Kannada
Fluent
Hindi
Advanced

Work Experience

Data Annotator & Generative AI Specialist at Innodata Inc.
January 1, 2026 - May 1, 2026
Processed and structured high-quality training datasets by classifying and tagging text samples across multiple annotation schemas to support AI/ML model development. Evaluated, rated, and compared AI-generated content across text, images, audio, and video for Generative AI workflows, assessing accuracy, reasoning, clarity, and overall quality against defined rubrics. Conducted bilingual content evaluation and rating tasks in English, Kannada, and Hinglish, contributing to multilingual dataset development and cross-lingual model performance. Reviewed annotation outputs against project guidelines, providing structured feedback to improve AI model performance and flagging inconsistencies before submission to the model training pipeline.
AI Prompt Designer & Data Annotator at OneForma (Appen)
February 1, 2025 - December 1, 2025
Executed NLP-focused data annotation and labeling tasks — including entity recognition, intent classification, and sentiment tagging — to build training datasets for large-scale machine learning pipelines across global clients. Designed and optimized prompt engineering strategies to enhance AI model reasoning, response accuracy, clarity, and natural language generation quality. Conducted content moderation and fact-checking across text, image, and video content for Meta platforms, assessing accuracy, relevance, and coherence, and validating information against credible external sources to ensure data integrity and compliance. Delivered QA reviews on transcription and translation datasets in English, Hindi, and Hinglish, verifying linguistic accuracy, label consistency, and adherence to annotation guidelines. Curated and submitted annotated microtask datasets meeting platform-specific quality thresholds, consistently achieving task approval rates above 95%.
Data Analysis Intern at Codec Technologies
February 1, 2024 - May 1, 2024
Categorized and tagged 5,000+ data points across structured and unstructured text datasets, generating clean training data for NLP and machine learning model pipelines. Developed Python scripts (Pandas, Seaborn) to preprocess raw datasets, eliminate noise, and visualize data distributions — reducing annotation team's preprocessing time by 25%. Validated annotation outputs for inter-annotator consistency and provided structured feedback on labelling accuracy, maintaining label accuracy above 95% before dataset handoff to the model training team. Drafted annotation guidelines and documented edge cases to standardize labelling processes, improving team consistency and judgment across projects.

Education

B.Tech in Information Technology at CMR University, Bengaluru
August 1, 2022 - May 1, 2026

Qualifications

Oracle Cloud Infrastructure: Data Science Professional
January 11, 2030 - July 5, 2026
Data Analysis using Excel
January 11, 2030 - July 5, 2026

Industry Experience

Software & Internet, Media & Entertainment, Professional Services