Hi, I'm Rishita Kalidindi, a highly skilled Data Engineer with over five years of experience building scalable data infrastructure and analytics solutions across cloud platforms. I specialize in developing ETL pipelines, optimizing database queries, and creating intuitive dashboards to deliver real-time insights for various business needs. I enjoy collaborating with cross-functional teams to enable data-driven decision-making and streamline workflows. My expertise spans from data warehousing in Snowflake and Azure to leveraging big data tools like Apache Spark and Hadoop. I'm passionate about turning complex data into actionable insights and continuously improving data processes.

Rishita Kalidindi

Hi, I'm Rishita Kalidindi, a highly skilled Data Engineer with over five years of experience building scalable data infrastructure and analytics solutions across cloud platforms. I specialize in developing ETL pipelines, optimizing database queries, and creating intuitive dashboards to deliver real-time insights for various business needs. I enjoy collaborating with cross-functional teams to enable data-driven decision-making and streamline workflows. My expertise spans from data warehousing in Snowflake and Azure to leveraging big data tools like Apache Spark and Hadoop. I'm passionate about turning complex data into actionable insights and continuously improving data processes.

Available to hire

Hi, I’m Rishita Kalidindi, a highly skilled Data Engineer with over five years of experience building scalable data infrastructure and analytics solutions across cloud platforms. I specialize in developing ETL pipelines, optimizing database queries, and creating intuitive dashboards to deliver real-time insights for various business needs.

I enjoy collaborating with cross-functional teams to enable data-driven decision-making and streamline workflows. My expertise spans from data warehousing in Snowflake and Azure to leveraging big data tools like Apache Spark and Hadoop. I’m passionate about turning complex data into actionable insights and continuously improving data processes.

See more

Experience Level

Expert
Expert
Expert
Expert
Expert
Expert
Expert
Expert
Intermediate
Intermediate
See more

Work Experience

Software Engineer / Data Engineer at McKinsey & Company
August 1, 2024 - June 23, 2024
Developed automated ETL pipelines using Python and SQL to enhance financial transaction processing, reducing manual effort by 140 hours per month and boosting data refresh speeds. Optimized SQL Server queries to lower latency for large datasets. Architected and maintained Azure infrastructure components to streamline financial data workflows. Created financial data marts powering dashboards for credit risk and fraud analytics. Automated ingestion of banking logs and claims data using MapReduce and custom Python scripts. Designed dimensional data models to support credit scoring for faster KPI access. Developed data validation frameworks to ensure ETL integrity. Integrated Snowflake for scalable data warehousing to accelerate reporting. Led migration of finance systems to Azure with full data integrity. Enhanced team-wide Git documentation and versioning, improving code reusability across 20+ engineers.
Software Engineer / Data Engineer at McKinsey & Company
August 1, 2024 - June 22, 2024
Developed automated ETL pipelines using Python and SQL to process financial transactions, reducing manual effort by 140 hours/month and improving data refresh speeds. Optimized SQL Server queries and indexing strategies, lowering query latency by 2 seconds for multi-million row datasets. Architected and maintained Azure infrastructure including Data Factory, Blob Storage, Data Lake, and Databricks, streamlining sensitive financial data workflows. Built financial data marts powering daily dashboards for stakeholders in credit risk, fraud analytics, and compliance. Automated ingestion of structured and unstructured datasets using MapReduce, HDFS, and custom Python scripts, handling 1TB+ of banking logs and claims data. Designed dimensional data models and star schemas to support credit scoring, enabling faster access to KPIs and accurate monthly metrics. Developed data validation frameworks using Python and SQL, ensuring high-integrity ETL with minimal compliance risk. Provisioned analyt
Graduate Research Assistant at University Of South Florida
March 1, 2023 - May 31, 2024
Designed and implemented ETL pipelines using SSIS to increase data accessibility by 60% for research analytics. Managed 20+ structured and semi-structured datasets related to child development. Automated data cleaning and transformation workflows, cutting processing time by 30%. Created 30+ interactive dashboards with Tableau, supporting faculty-led research. Conducted statistical analysis to ensure research accuracy. Collaborated with cross-functional teams to refine data collection and comply with governance standards.
Graduate Research Assistant at University Of South Florida
March 1, 2023 - May 31, 2024
Designed and implemented ETL pipelines using SSIS to extract, transform, and load data from blob storage and CSV files into Microsoft SQL Server, increasing data accessibility by 60% for research analytics. Managed and curated 20+ structured and semi-structured datasets related to child development assessments, leveraging advanced Excel functions and Python scripting for preprocessing and exploratory analysis. Automated data cleaning and transformation workflows using Pandas, NumPy, and SciPy, cutting data processing time by 30% and enabling quicker turnaround for insights generation. Built and delivered 30+ interactive dashboards and visualizations using Tableau, supporting faculty-led research by highlighting key behavioral trends and developmental outcomes. Conducted statistical analysis and data validation to ensure research integrity and accuracy, supporting evidence-based decision-making across 10+ academic reports. Collaborated with cross-functional research teams to refine data
ETL Developer / Data Engineer at Syniti
July 1, 2020 - August 31, 2022
Developed scalable ETL pipelines using Python, SQL, and Talend for large daily data transformations. Built modular workflows with Apache Airflow, reducing manual work. Designed and optimized data warehouses supporting marketing and finance analytics. Integrated real-time pipelines using Kafka and Spark. Built data lake architecture with Amazon S3 and Delta Lake for scalable storage. Applied data cleansing techniques with Pandas and NumPy. Created dashboards in Power BI and Excel. Automated metadata-driven pipeline configurations to reduce errors. Supported cloud development using AWS Glue, Lambda, and EC2. Participated in Agile sprints and maintained development with GitHub.
ETL Developer / Data Engineer at Syniti
July 1, 2020 - August 31, 2022
Developed scalable ETL pipelines using Python, SQL, and Talend, transforming data from multiple structured and semi-structured sources totaling over 20 GB/day. Built modular data workflows with Apache Airflow, reducing manual interventions and automating job dependencies. Designed and optimized data warehouses in SQL Server, enhancing query efficiency and supporting analytics for marketing and finance teams. Integrated real-time pipelines using Apache Kafka and Apache Spark, improving data availability across downstream applications. Built data lake architecture using Amazon S3 and Delta Lake, improving storage scalability and query flexibility. Applied Pandas and NumPy to cleanse inconsistent records across financial and customer datasets. Created Power BI dashboards and Excel-based reports, driving visibility for operations and leadership metrics. Automated metadata-driven pipeline configurations, enhancing reusability and reducing manual errors in transformations. Facilitated cloud-
Data Analyst at Cipla
August 1, 2019 - June 30, 2020
Analyzed pharmaceutical transaction data using Python to detect anomalies. Queried substance records via MySQL to support regulatory reporting. Built Power BI dashboards for real-time fraud metrics across supply chains. Automated ETL using Power Query and DAX, reducing manual report time by 16 hours weekly. Modeled refill rates and distribution anomalies in Excel for executive summaries. Collaborated with clinical teams to flag non-compliance, improving investigation accuracy. Managed version control with Git for reproducible pipelines. Delivered statistical findings through weekly dashboards aligned with compliance protocols.
Data Analyst at Cipla
August 1, 2019 - June 30, 2020
Analyzed over 500,000 pharmaceutical transactions using Python (Pandas, NumPy) to identify anomalies in prescription and distribution data. Queried controlled substance records using MySQL, extracting key insights across 20+ attributes to support regulatory reporting workflows. Built 10+ dynamic dashboards in Power BI showcasing real-time fraud detection metrics across regional supply chains. Automated ETL logic using Power Query and DAX, reducing manual report generation time by 16 hours per week. Modeled refill rates and distribution anomalies in Excel using advanced formulas and pivot tables for executive-level summaries. Collaborated with clinical and audit teams to flag non-compliant pharmacy behavior across 4 therapeutic categories, improving investigation accuracy. Managed 5 version-controlled repositories via Git, ensuring reproducible data pipelines and traceable audit workflows. Translated statistical findings into weekly visual summaries and dashboards aligned with pharma co

Education

Master’s of Science at University of South Florida
August 1, 2022 - May 31, 2024
Bachelor’s of Technology at BV Raju Institute of Technology
August 1, 2016 - June 30, 2020
Master’s of Science at University of South Florida
August 1, 2022 - May 31, 2024
Bachelor’s of Technology at BV Raju Institute of Technology
August 1, 2016 - June 30, 2020

Qualifications

Add your qualifications or awards here.

Industry Experience

Financial Services, Education, Healthcare, Professional Services, Software & Internet

Experience Level

Expert
Expert
Expert
Expert
Expert
Expert
Expert
Expert
Intermediate
Intermediate
See more