I'm Vinay Kumar Goud, an AI engineer specializing in backend systems and ML platforms with 5+ years of experience building scalable APIs, asynchronous workers, model-serving workflows, and cloud-based ML infrastructure. I enjoy turning research prototypes into reliable production pipelines, collaborating with product managers, researchers, and infra teams to improve reliability, throughput, and deployment efficiency, and continually learning to apply responsible AI and sound system design practices.

Vinay Kumar Goud

I'm Vinay Kumar Goud, an AI engineer specializing in backend systems and ML platforms with 5+ years of experience building scalable APIs, asynchronous workers, model-serving workflows, and cloud-based ML infrastructure. I enjoy turning research prototypes into reliable production pipelines, collaborating with product managers, researchers, and infra teams to improve reliability, throughput, and deployment efficiency, and continually learning to apply responsible AI and sound system design practices.

Available to hire

I’m Vinay Kumar Goud, an AI engineer specializing in backend systems and ML platforms with 5+ years of experience building scalable APIs, asynchronous workers, model-serving workflows, and cloud-based ML infrastructure.

I enjoy turning research prototypes into reliable production pipelines, collaborating with product managers, researchers, and infra teams to improve reliability, throughput, and deployment efficiency, and continually learning to apply responsible AI and sound system design practices.

See more

Experience Level

Expert
Expert
Expert
Expert
Expert
Expert
Expert
Expert
See more

Language

Bashkir
Advanced

Work Experience

Machine Learning Engineer at Scale AI
May 1, 2025 - Present
Designed and built Python services for scheduling LLM evaluation jobs, collecting responses, storing benchmark results, and tracking regressions across releases. Implemented 1M+ evaluation runs/week across 200+ GPU nodes on Kubernetes; improved regression detection time by 45% with automated pipelines. Architected queue-based orchestration with retry logic enabling concurrent processing of 500K+ LLM responses weekly and graceful failure handling. Developed scoring and review APIs for LLM-as-judge workflows with rubrics, human review loops, audit trails, and structured feedback. Added observability with structured logging, metrics, dashboards, alerts, and runbooks; collaborated to convert research prototypes into repeatable evaluation workflows and release guidelines.
Software Engineer — Machine Learning at NVIDIA
May 1, 2024 - April 1, 2025
Built multimodal data ingestion, preprocessing pipelines, and PyTorch data loaders for video, depth, segmentation, LiDAR, and simulation metadata; improved GPU training throughput by 35-40% through batch size tuning, enabling mixed precision, reducing I/O bottlenecks, and profiling pipeline stages. Developed experiment tracking, model versioning, CI/CD automation, and containerized training workflows using Docker, Kubernetes, and MLflow. Supported GPU-backed inference deployment by profiling request latency, memory usage, batching, and failure modes, and helped optimize TensorRT/CUDA pipelines. Built AWS workflows using S3, EKS, SageMaker, and Lambda for scalable data preprocessing, model training, and evaluation. Validated dataset integrity, schema consistency, missing labels, and corrupted simulation artifacts before training; wrote runbooks for dataset handling and training incident response.
Software Engineer at Accenture
December 1, 2020 - June 1, 2023
Built backend services and ML deployment workflows for enterprise risk and fraud analytics applications, including model inference APIs, feature stores, and prediction logging. Standardized deployment pipelines for 80+ models across 20 internal use cases, reducing manual effort by 45% through reusable CI/CD templates and structured monitoring. Implemented drift detection, data quality checks, and dashboard alerts to catch production issues early and reduce incident frequency by 30%. Developed infrastructure on AWS SageMaker, EKS, S3, and Lambda for batch and real-time predictions, including containerization and automated rollback. Wrote runbooks for deployment, rollback, environment configuration, and validation checks; collaborated with data engineers, product managers, and compliance teams to implement explainability, audit logging, and regulatory controls for models in production.

Education

M.S. in Computer Science at University of Central Missouri
January 11, 2030 - June 29, 2026

Qualifications

Oracle MySQL Implementation Specialist
January 11, 2030 - June 29, 2026
AWS Solutions Architecture Job Simulation
January 11, 2030 - June 29, 2026

Industry Experience

Software & Internet, Professional Services, Media & Entertainment