Available to hire
Senior Data Scientist with ~5 years of experience specializing in machine learning, NLP, and Generative AI. Builds scalable RAG pipelines and LLM-powered/agentic AI systems, with a focus on production-ready model development, deployment, evaluation, and reliability.
Experienced across finance, healthcare, and retail domains, leveraging Python, PyTorch/TensorFlow, SQL, PySpark, and vector search (FAISS) alongside MLOps tooling such as AWS, Docker, Kubernetes, and MLflow. Strong in hallucination mitigation, grounding strategies, prompt/versioning experimentation, and end-to-end ETL-to-deployment pipelines.
Skills
Experience Level
Expert
Expert
Expert
Expert
Expert
Expert
Expert
Expert
Expert
Expert
Expert
Expert
Intermediate
Intermediate
Work Experience
Sr. Data Scientist at JP Morgan Chase & Co
January 1, 2024 - PresentDesigned and implemented a retrieval-augmented generation (RAG) pipeline for contextual summarization of financial documents (earnings reports, analyst notes, and market news), reducing manual research effort by 30% for internal analyst workflows. Built semantic chunking aligned to financial document structure (300–500 tokens with 10–15% overlap), improving retrieval relevance by 22%. Generated dense embeddings using sentence-transformer models in PyTorch and indexed them in FAISS, enabling low-latency top-k retrieval (<120ms) and increasing retrieval accuracy by 25% over keyword baselines. Developed LLM-based summarization and Q&A pipelines using transformer architectures with structured prompts. Implemented hallucination mitigation by grounding outputs in retrieved context using controlled decoding and fallback responses, reducing unsupported outputs by 35%. Created a retrieval evaluation framework (Precision@k + human validation), set up prompt versioning/A-B testing, and improv
Data Scientist at Dixon Technology
May 1, 2020 - July 1, 2022Performed exploratory data analysis using Pandas, NumPy, and SQL on structured enterprise datasets (300K+ records), identifying data quality issues, correlations, and feature patterns to improve downstream modeling and data reliability. Built and evaluated ML models (Linear/Logistic Regression, Decision Trees, XGBoost) for churn and risk prediction, achieving 82–85% validation accuracy. Engineered features from transactional and behavioral data using aggregation, encoding, and scaling, improving performance and stability. Applied interpretability using SHAP/feature importance to support stakeholder understanding. Developed healthcare forecasting models (XGBoost, LightGBM, LSTM) on patient/epidemiological datasets (400K+ records), improving forecast accuracy by 12% and reducing readmission risk for high-risk groups. Automated ETL and MLOps pipelines using PySpark, Airflow, MLflow, Docker, and Kubernetes, processing 700K–900K records daily and reducing data processing time by 35%. Bu
Education
Master of Science in Computer Science at University at Albany, New York
May 1, 2024 - May 1, 2024Qualifications
Industry Experience
Financial Services, Healthcare, Retail
Skills
Experience Level
Expert
Expert
Expert
Expert
Expert
Expert
Expert
Expert
Expert
Expert
Expert
Expert
Intermediate
Intermediate
Hire a Data Scientist
We have the best data scientist experts on Twine. Hire a data scientist in Albany today.