I’m an AI/ML Engineer who enjoys turning messy, high-volume data into reliable production models that solve real business problems. Over the past few years, I’ve built and deployed end-to-end machine learning systems—ranging from classification and predictive models to real-time pipelines—while tracking performance with clear metrics and making sure results translate into measurable outcomes like reduced latency, improved accuracy, and stronger revenue impact. I also work hands-on with modern GenAI patterns, including Retrieval-Augmented Generation (RAG), prompt/agentic tool use, and fine-tuning workflows, with a strong focus on deployment, monitoring, and governance. From containerized inference on cloud infrastructure to streaming feature engineering and secure document retrieval, I like building systems that are scalable, observable, and useful for teams and users in production.

Azmaan Amin Hemraj

I’m an AI/ML Engineer who enjoys turning messy, high-volume data into reliable production models that solve real business problems. Over the past few years, I’ve built and deployed end-to-end machine learning systems—ranging from classification and predictive models to real-time pipelines—while tracking performance with clear metrics and making sure results translate into measurable outcomes like reduced latency, improved accuracy, and stronger revenue impact. I also work hands-on with modern GenAI patterns, including Retrieval-Augmented Generation (RAG), prompt/agentic tool use, and fine-tuning workflows, with a strong focus on deployment, monitoring, and governance. From containerized inference on cloud infrastructure to streaming feature engineering and secure document retrieval, I like building systems that are scalable, observable, and useful for teams and users in production.

Available to hire

I’m an AI/ML Engineer who enjoys turning messy, high-volume data into reliable production models that solve real business problems. Over the past few years, I’ve built and deployed end-to-end machine learning systems—ranging from classification and predictive models to real-time pipelines—while tracking performance with clear metrics and making sure results translate into measurable outcomes like reduced latency, improved accuracy, and stronger revenue impact.

I also work hands-on with modern GenAI patterns, including Retrieval-Augmented Generation (RAG), prompt/agentic tool use, and fine-tuning workflows, with a strong focus on deployment, monitoring, and governance. From containerized inference on cloud infrastructure to streaming feature engineering and secure document retrieval, I like building systems that are scalable, observable, and useful for teams and users in production.

See more

Experience Level

Expert
Expert
Expert
Expert
Expert
Expert
Expert
Expert
Expert
Expert
Expert
Intermediate
Intermediate
See more

Language

Work Experience

AI/ML Engineer at BNY, Texas
October 1, 2025 - Present
Optimized the model inference pipeline by containerizing deployment workflows with Docker and deploying on AWS SageMaker; used CloudWatch for production monitoring and reduced average response latency by 35%. Automated model retraining workflows by integrating GitHub Actions with Airflow-orchestrated pipelines and MLflow tracking, cutting redeployment cycle time from two weeks to under 4 hours. Built a RAG pipeline with LangChain and a self-hosted vector database (Milvus) to enable secure internal document search for 50+ users with data governance. Fine-tuned a domain-specific language model using LoRA, validated performance via tracked accuracy and F1 metrics, and deployed it via a containerized API with a 12% accuracy improvement. Designed real-time feature engineering pipelines using Apache Spark and Kafka for streaming transaction data to support downstream detection models processing 1M+ records daily with improved data quality.
AI/ML Engineer at BNY
October 1, 2025 - Present
Optimized production ML inference by containerizing deployment workflows with Docker on AWS SageMaker and monitoring via CloudWatch, reducing average response latency by 35%. Automated model retraining using GitHub Actions with Airflow-orchestrated pipelines and MLflow tracking, cutting redeployment time from two weeks to under 4 hours. Built a RAG pipeline using LangChain with a self-hosted vector database (Milvus) for secure internal document search for 50+ users with governance controls. Fine-tuned a domain-specific LLM with LoRA, validated metrics (accuracy and F1), and deployed via containerized APIs to improve accuracy by 12%. Designed real-time feature engineering pipelines with Apache Spark and Kafka to stream transaction data supporting downstream detection models processing 1M+ records daily with improved data quality.
ML Engineer at Vivma Software Inc, India
August 1, 2022 - June 30, 2024
Built a collaborative filtering recommendation engine in Python and TensorFlow, preceded by SQL-based customer segmentation on 1M+ transactions, driving a 24% improvement in conversion rates. Deployed XGBoost and scikit-learn real-time pricing models processing 50K+ products hourly; supported demand analysis in Power BI and delivered 16% annual revenue growth for a retail client. Deployed ML models on AWS SageMaker with Docker containerization to enable automated training and production inference, with Tableau dashboards tracking accuracy drift and providing performance insights to stakeholders. Performed EDA, statistical profiling, and Pandas/NumPy feature engineering on 2M+ customer records to identify outliers and high-correlation variables, improving model performance metrics by 19%. Integrated ML models into AWS microservices via REST APIs for real-time inference and ongoing performance analysis tied to business-impact reporting.
ML Engineer at Vivma Software Inc
August 1, 2022 - June 1, 2024
Built a collaborative filtering recommendation engine using Python and TensorFlow, preceded by SQL-based customer segmentation on 1M+ transactions, improving conversion rates by 24%. Deployed an XGBoost/scikit-learn real-time pricing model processing 50K+ products hourly and supported analytics in Power BI for demand insights, contributing to 16% annual revenue growth. Deployed ML models on AWS SageMaker using Docker-based training/inference workflows and used Tableau dashboards to monitor accuracy drift and share performance insights. Performed EDA, statistical profiling, and Pandas/NumPy feature engineering on 2M+ customer records, identifying outliers and key correlated features to improve model performance by 19%. Integrated models into AWS microservices via REST APIs for real-time inference, maintaining performance analysis and business-impact reporting.

Education

Master of Science in Computer Science at University of North Texas
January 11, 2030 - May 1, 2026
Bachelor of Engineering in Computer Science at Osmania University
January 11, 2030 - July 1, 2024
Master of Science in Computer Science at University of North Texas
January 1, 2024 - May 1, 2026
Bachelor of Engineering in Computer Science at Osmania University
January 1, 2020 - July 1, 2024
Master of Science in Computer Science at University of North Texas
January 11, 2030 - May 1, 2026
Bachelor of Engineering in Computer Science at Osmania University
July 1, 2022 - July 1, 2024
Master of Science in Computer Science at University of North Texas
January 11, 2030 - May 1, 2026
Bachelor of Engineering in Computer Science at Osmania University
January 1, 2024 - July 1, 2024

Qualifications

Add your qualifications or awards here.

Industry Experience

Financial Services, Software & Internet, Computers & Electronics, Education, Professional Services