Available to hire
I’m Shubham Shekhar Patil, a Data Scientist who designs and deploys production-grade ML, NLP, and Generative AI solutions across enterprise environments. I enjoy turning complex data into practical, scalable systems that drive measurable business impact.
I translate business problems into data-driven strategies, using statistical modeling, experimentation, and robust MLOps to deliver reliable predictions and real-time insights. My hands-on experience spans Python, SQL, cloud platforms, and end-to-end deployment with FastAPI, Docker, Kubernetes, and MLflow.
Skills
Work Experience
Data Scientist at HCL Tech
January 1, 2025 - PresentDesigned and deployed a retrieval-augmented generation (RAG) based knowledge system to boost internal knowledge retrieval accuracy by 38%. Built a LangChain-ChromaDB pipeline and fine-tuned transformer models for enterprise knowledge management. Implemented a LangGraph workflow and MCP integrations, and rolled out n8n automation to enable business teams—cutting document turnaround on 10K+ monthly requests to under 15 minutes. Strengthened financial forecasting by applying regression analysis, hypothesis testing, and probabilistic methods in Python and R. Enhanced entity extraction across unstructured documents by fine-tuning BERT models with PyTorch and Hugging Face for classification and NER. Reduced real-time inference latency by 40% via ONNX runtime, deployed through FastAPI on Docker and Kubernetes for scalable cloud serving. Integrated ML outputs with SAP ERP and Oracle ERP workflows for forecasting, exception monitoring, and business planning.
Data Scientist at Mindtree
August 1, 2021 - July 1, 2023Expanded campaign targeting by developing propensity and segmentation models using Python, SQL, Scikit-learn, and XGBoost across multi-source datasets, increasing targeting effectiveness by 27%. Reduced batch data processing time by 35% by engineering distributed pipelines with PySpark, Apache Spark, and Kafka for large-scale transactional and behavioral data. Improved forecasting across seasonal cycles by applying statistical modeling, probability analysis, and hypothesis testing using R and SAS for demand and performance analytics. Implemented LSTM-based forecasting and anomaly detection in TensorFlow for supply and demand monitoring. Shortened model release cycles from weeks to days by establishing CI/CD with MLflow, Git, and Jenkins. Built KPI dashboards in Tableau and Power BI to unify analytics for marketing, operations, and leadership. Reduced cloud training costs by 18% by migrating to spot instances and right-sizing Databricks clusters on AWS.
Data Analyst at Mindtree
June 1, 2020 - July 1, 2021Reduced manual reporting effort by 40% by designing SQL- and Python-based reporting pipelines for finance and operations. Improved data quality and automated reconciliation across enterprise datasets with Pandas and NumPy. Cut KPI report generation time by 30% by building interactive dashboards in Power BI and Excel. Eliminated 60% of reconciliation issues across Oracle data flows by redesigning ETL workflows with SQL and Informatica. Identified a 14% lift in customer engagement through A/B testing on HubSpot campaigns, informing rollout decisions.
Education
M.S. in Information Systems at Syracuse University
January 11, 2030 - May 1, 2025Certificate of Advanced Study in Data Science at Syracuse University
January 11, 2030 - May 1, 2025B.E. in Computer Engineering at University of Mumbai
January 11, 2030 - October 1, 2020Qualifications
Microsoft Power BI
January 11, 2030 - June 29, 2026Microsoft Azure Data Engineer
January 11, 2030 - June 29, 2026Microsoft Excel
January 11, 2030 - June 29, 2026OCI Data Science
January 11, 2030 - June 29, 2026OCI Generative AI
January 11, 2030 - June 29, 2026Industry Experience
Software & Internet, Professional Services, Financial Services, Education, Healthcare
Skills
Hire a Data Scientist
We have the best data scientist experts on Twine. Hire a data scientist today.