Hi, I’m Rohith Arepally, a data engineer with 4+ years of experience designing and building scalable cloud-based data platforms across healthcare, banking, and enterprise domains using Azure and AWS. I specialize in end-to-end ELT/ETL pipelines, data lakes, data warehouses, real-time streaming, and analytics platforms with Databricks, Azure Synapse Analytics, Snowflake, AWS Glue, PySpark, and Apache Airflow. I’m also exploring AI/GenAI enablement through RAG architectures, LangChain, and OpenAI APIs to deliver analytics-ready datasets for BI, ML, and enterprise reporting. I focus on optimizing data processing, boosting platform performance, and supporting large-scale workloads in Agile environments. I collaborate with data scientists and ML engineers, implement data governance and RBAC, and build reusable CI/CD pipelines to accelerate deployment and ensure quality. I enjoy turning complex data into actionable insights that empower business decisions and self-service analytics.

Rohith Arepally

Hi, I’m Rohith Arepally, a data engineer with 4+ years of experience designing and building scalable cloud-based data platforms across healthcare, banking, and enterprise domains using Azure and AWS. I specialize in end-to-end ELT/ETL pipelines, data lakes, data warehouses, real-time streaming, and analytics platforms with Databricks, Azure Synapse Analytics, Snowflake, AWS Glue, PySpark, and Apache Airflow. I’m also exploring AI/GenAI enablement through RAG architectures, LangChain, and OpenAI APIs to deliver analytics-ready datasets for BI, ML, and enterprise reporting. I focus on optimizing data processing, boosting platform performance, and supporting large-scale workloads in Agile environments. I collaborate with data scientists and ML engineers, implement data governance and RBAC, and build reusable CI/CD pipelines to accelerate deployment and ensure quality. I enjoy turning complex data into actionable insights that empower business decisions and self-service analytics.

Available to hire

Hi, I’m Rohith Arepally, a data engineer with 4+ years of experience designing and building scalable cloud-based data platforms across healthcare, banking, and enterprise domains using Azure and AWS. I specialize in end-to-end ELT/ETL pipelines, data lakes, data warehouses, real-time streaming, and analytics platforms with Databricks, Azure Synapse Analytics, Snowflake, AWS Glue, PySpark, and Apache Airflow. I’m also exploring AI/GenAI enablement through RAG architectures, LangChain, and OpenAI APIs to deliver analytics-ready datasets for BI, ML, and enterprise reporting.

I focus on optimizing data processing, boosting platform performance, and supporting large-scale workloads in Agile environments. I collaborate with data scientists and ML engineers, implement data governance and RBAC, and build reusable CI/CD pipelines to accelerate deployment and ensure quality. I enjoy turning complex data into actionable insights that empower business decisions and self-service analytics.

See more

Experience Level

Expert
Expert
Expert
Expert
Expert
Expert
Expert
Expert
Expert
Expert
Intermediate
See more

Work Experience

Data Engineer at Adobe
January 1, 2025 - Present
Designed and built scalable ELT pipelines using Azure Data Factory, Databricks, PySpark, and Azure Synapse Analytics; implemented Medallion Architecture on Delta Lake with ADLS Gen2; developed dimensional data models and data marts in Snowflake; established data quality, schema validation, and monitoring; built real-time streaming with Azure Event Hub and Databricks Structured Streaming; orchestrated end-to-end workflows with ADF and Apache Airflow; implemented data governance with Unity Catalog and RBAC; built reusable Python ETL frameworks; published curated datasets and semantic models for self-service analytics; developed RAG pipelines with LangChain and OpenAI APIs; optimized dbt models and Delta Lake workloads; collaborated with data scientists for feature engineering; set up CI/CD with GitHub Actions and Azure DevOps; optimized workloads with partitioning and caching to reduce cloud costs.
Data Engineer at Wells Fargo
June 1, 2020 - July 1, 2023
Designed scalable data pipelines using AWS Glue, PySpark, and Python processing 10M+ banking transactions daily; built centralized data lake on S3 with Lake Formation and Glue Data Catalog; used DynamoDB for high-volume, low-latency access; developed ETL workflows with Glue and Apache Airflow; implemented real-time streaming with Kinesis Data Streams and Lambda; developed dimensional data models and optimized SQL workloads in Redshift for reporting and risk analytics; automated data validation and reconciliation with Python; configured CloudWatch dashboards and alerts for pipeline health; migrated legacy ETL to AWS cloud, reducing costs and improving scalability; built Power BI dashboards for financial reporting and regulatory compliance; implemented security controls including encryption, IAM policies, and RBAC.

Education

MS in Data Science at University of Texas at Arlington
August 1, 2023 - May 1, 2025
MS in Data Science at University of Texas at Arlington
August 1, 2023 - May 1, 2025

Qualifications

Add your qualifications or awards here.

Industry Experience

Healthcare, Financial Services, Professional Services, Software & Internet