I am a data scientist with 4+ years of experience building and deploying machine learning, deep learning, and AI-powered solutions for predictive analytics, NLP, computer vision, and Generative AI applications. I’m proficient in Python, SQL, and ML/LLM frameworks (PyTorch, TensorFlow, Scikit-Learn, LangChain, Hugging Face) with strong capabilities in feature engineering, model optimization, prompt engineering, RAG workflows, and advanced analytics to solve complex business problems. I have end-to-end experience across EDA, big data processing (Hadoop, Spark), cloud deployment (AWS, Azure), and visualization (Tableau, Power BI), delivering actionable insights and scalable solutions that drive business impact.

Joshithavelkur

I am a data scientist with 4+ years of experience building and deploying machine learning, deep learning, and AI-powered solutions for predictive analytics, NLP, computer vision, and Generative AI applications. I’m proficient in Python, SQL, and ML/LLM frameworks (PyTorch, TensorFlow, Scikit-Learn, LangChain, Hugging Face) with strong capabilities in feature engineering, model optimization, prompt engineering, RAG workflows, and advanced analytics to solve complex business problems. I have end-to-end experience across EDA, big data processing (Hadoop, Spark), cloud deployment (AWS, Azure), and visualization (Tableau, Power BI), delivering actionable insights and scalable solutions that drive business impact.

Available to hire

I am a data scientist with 4+ years of experience building and deploying machine learning, deep learning, and AI-powered solutions for predictive analytics, NLP, computer vision, and Generative AI applications. I’m proficient in Python, SQL, and ML/LLM frameworks (PyTorch, TensorFlow, Scikit-Learn, LangChain, Hugging Face) with strong capabilities in feature engineering, model optimization, prompt engineering, RAG workflows, and advanced analytics to solve complex business problems.

I have end-to-end experience across EDA, big data processing (Hadoop, Spark), cloud deployment (AWS, Azure), and visualization (Tableau, Power BI), delivering actionable insights and scalable solutions that drive business impact.

See more

Experience Level

Expert
Expert
Expert
Expert
Expert
Expert
Expert
Expert
Expert
Expert
Expert
Expert
Expert
Expert
Expert
Intermediate
See more

Language

English
Fluent

Work Experience

Machine Learning Research Assistant at DMLAB@GSU
August 1, 2025 - Present
Designed and deployed GenAI-powered research knowledge assistant to help researchers access project documentation, research papers, experiment logs, datasets, and model outputs, reducing manual search time. Built a Retrieval-Augmented Generation (RAG) pipeline using FAISS/Pinecone, embeddings, semantic retrieval, document chunking, and context engineering to improve answer relevance and enable accurate contextual Q&A across research artifacts. Developed AI agents and multi-agent workflows for research paper summarization, document analysis, dataset validation, experiment tracking, and automated decision support, reducing repetitive manual review effort. Engineered MCP servers and custom tool integrations to securely connect LLM agents with research datasets, APIs, vector databases, experiment-tracking tools, and project-specific knowledge sources, improving access to research artifacts. Authored reusable prompt templates, evaluation harnesses, slash commands, and agent orchestration wo
Data Science Intern at ASTM International
May 1, 2025 - August 1, 2025
Architected and deployed a data extraction pipeline to convert unstructured PDF and Excel reports into structured datasets, reducing manual data preparation by 75%. Processed and analyzed over 120 materials test reports using document AI and table extraction tools including pdfplumber, Camelot, PyMuPDF, pytesseract, Table Transformer, LayoutParser, and MinerU, achieving 85% accuracy in table reconstruction and improving structural consistency by 40% across complex scientific formats. Integrated multimodal LLMs and RAG into the document parsing workflow with text cleaning, tokenization, embedding normalization, and optimized chunking, improving data extraction efficiency and retrieval accuracy by 25%. Collaborated with materials science and quality assurance teams to define validation rules for key metrics such as modulus and tensile strength, reducing reporting inconsistencies by 50%. Applied LLMs, computer vision, and Python-based automation to solve complex scientific data extraction
Research Assistant – AI/ML Engineering at Georgia State University
August 1, 2024 - May 1, 2025
Generated high-fidelity speckle imaging datasets using HCIPy with multi-layer atmospheric turbulence models, Zernike polynomials, realistic noise injection, and efficient multiprocessing for fully automated HDF5 dataset creation and management. Constructed CNN, ResNet, and FNO architectures for spatiotemporal wind prediction, leveraging local spatial feature extraction and global frequency-domain modeling. Incorporated frame differencing and sliding-window framing (8–64 frames) to capture dynamic motion patterns and preserve long-range temporal dependencies for accurate multi-layer wind speed and direction forecasting. Executed advanced loss functions including MSE and VICReg for self-supervised learning, enforcing variance preservation and invariance across augmentations. Established a scalable PyTorch pipeline with multiprocessing, GPU acceleration, Docker containers, and Kubernetes for reliable production execution, integrating WandB for hyperparameter tuning and metric visualizat
Senior Data Scientist at ExxonMobil (HCLTech)
December 1, 2023 - August 31, 2024
Built a deep learning-based demand forecasting model using LSTM, increasing forecasting accuracy by 15% and turnover by 7%. Analyzed customer churn with gradient boosting (XGBoost) and random forest, improving retention strategies and reducing churn by 12%. Applied NLP with transformer-based models (e.g., BERT) to analyze customer feedback and market reviews, generating actionable insights. Leveraged Spark and Hadoop for large-scale data processing, and deployed on AWS (S3, EC2, SageMaker) to enhance scalability and reduce deployment time by 30%. Created reusable Airflow ETL templates to standardize data workflows, improving processing speed by 35%.
Senior Data Scientist at ExxonMobil | HCLTech
December 1, 2023 - August 1, 2024
Built a deep learning-based demand forecasting model using LSTM networks, increasing forecasting accuracy by 15% and improving inventory management and turnover. Used ML to analyze customer churn with gradient boosting (XGBoost) and random forest classifiers, reducing churn by 12%. Applied NLP techniques including transformer-based models like BERT to analyze customer feedback and market reviews, providing actionable insights to enhance product quality and service offerings. Leveraged Apache Spark and Hadoop to process large-scale datasets, boosting processing speed by 35%. Deployed models on AWS (S3, EC2, SageMaker) to improve scalability and reduce deployment time by 30%. Created reusable Airflow ETL templates to standardize data workflows and accelerate project delivery by 40%.
Data Scientist at Dollar General (HCLTech)
July 1, 2022 - December 31, 2023
Extracted, cleaned, and analyzed 10M+ rows of transactional data using SQL and Python, identifying purchasing patterns to optimize inventory, reducing excess stock and stockouts by 15%. Performed customer segmentation with K-means and DBSCAN, guiding targeted marketing with a 25% ROI uplift. Applied time-series forecasting (ARIMA, exponential smoothing) for demand planning, reducing overstock costs by 18%. Conducted A/B testing to optimize promotional tactics, achieving a 10% uplift in promotional efficiency. Built dashboards in Tableau and Power BI for KPI tracking, and developed a Ticket Analysis Dashboard with predictive insights for real-time monitoring.
Data Scientist at Dollar General | HCLTech
July 1, 2022 - December 1, 2023
Extracted, cleaned, and analyzed over 10 million rows of transactional data using SQL and Python (Pandas, NumPy) to identify purchasing patterns for inventory optimization, reducing overstock and stockouts by 15%. Performed customer segmentation using K-means and DBSCAN to guide targeted marketing and increase campaign ROI by 25%. Applied time-series forecasting (ARIMA, exponential smoothing) to predict seasonal demand, lowering overstock costs by 18%. Conducted A/B testing on promotional campaigns, achieving a 10% uplift. Optimized SQL queries for faster data retrieval (up to 25%), and built dashboards in Tableau and Power BI to boost reporting efficiency by 20%. Automated reporting workflows with Excel VBA macros.

Education

Master's in Data Science and Analytics at Georgia State University
January 11, 2030 - June 30, 2026
B.Tech. in ECE at Sreenivasa Institute of Technology and Management Studies
January 11, 2030 - June 30, 2026
Master’s in Data Science and Analytics at Georgia State University
January 11, 2030 - June 30, 2026
B.Tech. in Electronics and Communication Engineering at Sreenivasa Institute of Technology and Management Studies
January 11, 2030 - June 30, 2026

Qualifications

Add your qualifications or awards here.

Industry Experience

Software & Internet, Professional Services, Energy & Utilities, Manufacturing, Education, Computers & Electronics