Staff AI/ML Engineer with 12 years of experience deploying production-grade machine learning systems across AWS and enterprise research environments. Deep expertise in Python and modern deep learning frameworks (PyTorch, TensorFlow), distributed training, and MLOps practices that keep AI systems reliable, performant, and governable. Leads end-to-end delivery from scalable data pipelines and model monitoring/versioning to automated retraining, inference optimization, and secure REST API integration. Strong collaborator across product and engineering teams, with a focus on SLAs, observability, and responsible GenAI/LLMOps implementations.

Rishav Chakravarti

Staff AI/ML Engineer with 12 years of experience deploying production-grade machine learning systems across AWS and enterprise research environments. Deep expertise in Python and modern deep learning frameworks (PyTorch, TensorFlow), distributed training, and MLOps practices that keep AI systems reliable, performant, and governable. Leads end-to-end delivery from scalable data pipelines and model monitoring/versioning to automated retraining, inference optimization, and secure REST API integration. Strong collaborator across product and engineering teams, with a focus on SLAs, observability, and responsible GenAI/LLMOps implementations.

Available to hire

Staff AI/ML Engineer with 12 years of experience deploying production-grade machine learning systems across AWS and enterprise research environments. Deep expertise in Python and modern deep learning frameworks (PyTorch, TensorFlow), distributed training, and MLOps practices that keep AI systems reliable, performant, and governable.

Leads end-to-end delivery from scalable data pipelines and model monitoring/versioning to automated retraining, inference optimization, and secure REST API integration. Strong collaborator across product and engineering teams, with a focus on SLAs, observability, and responsible GenAI/LLMOps implementations.

See more

Language

English
Advanced

Work Experience

Staff Applied AI Engineer at Scale AI
August 1, 2025 - Present
Architected and delivered enterprise GenAI platforms on AWS using governed LLMOps pipelines, secure RAG architectures, and centralized model registries. Designed agentic AI systems coordinating retrieval, tool invocation, and multi-step reasoning with human-in-the-loop gates while preserving audit trails. Containerized and orchestrated GPU-aware deep learning workloads on Kubernetes with Terraform and optimized inference serving (batched requests and quantized models), improving throughput and p99 latency for production API endpoints. Built traceable data pipelines for ingesting, validating, versioning, and evaluating generative datasets with reproducible LLMOps cycles and regression testing. Operationalized MLOps end-to-end with model versioning, shadow deployments, drift detection, automated rollback, and real-time monitoring/diagnostics, reducing failed releases and maintaining high availability. Developed secure REST APIs with FastAPI, PostgreSQL, and Redis including RBAC, rate lim
Senior Machine Learning Engineer at Amazon Web Services (AWS)
July 1, 2020 - August 1, 2025
Led machine learning engineering for customer-facing AWS AI services by building and deploying production-grade PyTorch/TensorFlow models on SageMaker and Kubernetes with CI/CD, model versioning, monitoring, and performance evaluation across regions. Designed scalable telemetry data pipelines using Python, SQL, and Apache Spark to support feature engineering, training, and real-time inference at global scale. Optimized inference with quantization, GPU batching, and TensorRT to reduce latency and serving infrastructure costs while maintaining accuracy and reliability via automated benchmark regression checks and load testing. Established MLOps foundations including automated retraining, experiment tracking, model registry, canary releases, monitoring dashboards, runbooks, and incident response procedures. Built REST APIs/back-end services with FastAPI and SQL integrating predictions with secure authentication, throttling, structured logging, and audit requirements. Drove technical desig
Machine Learning Engineer at IBM Research
February 1, 2015 - July 1, 2020
Developed deep learning models in Python and TensorFlow for enterprise research projects, including supervised/unsupervised pipelines handling large-scale structured and unstructured data (text, logs, sensor data). Built distributed training systems on Kubernetes and Docker, scaling experiments across shared GPU clusters with optimized data loading and parallelization, including fault-tolerant checkpointing and recovery. Designed and implemented data processing and feature engineering workflows using Python, SQL, and Apache Spark with reproducible transformations, dataset versioning, validation checks, and performance optimization. Deployed proof-of-concept ML services as REST APIs using Docker and Kubernetes with health checks/readiness probes and load testing. Implemented model evaluation/validation/versioning practices for auditable research records and automated evaluation reporting. Collaborated with researchers and engineers on technical design/code reviews and contributed to sha
Research Assistant at Event and Pattern Detection Lab
August 1, 2014 - December 31, 2014
Applied Python and ML to detect patterns and anomalies in large-scale time-series event logs, improving pattern detection accuracy and supporting downstream statistical analysis. Implemented event pattern detection algorithms (sliding window, clustering, statistical thresholding) and evaluated using precision/recall/F1 and confusion matrices with documented reproducible results for publications. Cleaned/aggregated structured and unstructured datasets with Python/SQL including missing data imputation, outlier removal, and normalization; produced visualizations and annotated data dictionaries to inform experimental design and hypothesis testing. Maintained code repositories, tagged dataset versions, and configuration files for reproducibility and transparent handoffs. Automated data cleaning and feature normalization pipelines with reusable, configuration-driven workflows and structured logging to reduce manual processing time. Presented research results and supported reproducibility rev
Software Engineering Intern at Knewton
July 1, 2014 - August 1, 2014
Developed and maintained backend services in Python for a personalized learning platform, implementing REST APIs and integrating with SQL databases for student performance tracking and recommendations. Built data transformation scripts to convert raw interaction logs into structured analytical metrics for dashboards and adaptive learning workflows with automated jobs and validation rules to catch missing/malformed records. Wrote unit and integration tests using pytest to improve coverage and reduce defects; contributed to CI improvements. Participated in code reviews and technical design discussions applying modularity, testing, and performance best practices. Automated recurring data extraction/reporting workflows to reduce manual effort and errors, and created parameterized report templates for multiple use cases. Documented API schemas, architecture, and onboarding guides for data engineering workflows to support faster team ramp-up.

Education

Master of Information Systems and Management at Carnegie Mellon University
January 1, 2013 - January 1, 2014
B. Sc., Computer Science at Bucknell University
January 1, 2006 - January 1, 2010

Qualifications

Add your qualifications or awards here.

Industry Experience

Software & Internet, Computers & Electronics, Professional Services