I am a results-driven Data Engineer with 8+ years of experience designing, optimizing, and automating data pipelines and data lakes in large-scale enterprise environments. I excel at building scalable batch and streaming data solutions using Snowflake, AWS, Python, dbt, PySpark, Apache Iceberg, Airflow, and AWS Glue to enable real-time analytics, BI, and reliable cloud data ecosystems. I have strong experience in data quality frameworks, governance, and performance tuning across AWS Redshift and Delta Lake, enabling observability, cost efficiency, and reliability while collaborating with analytics, BI, and ML teams to deliver trusted datasets for forecasting and decision-making.

VEDA C

I am a results-driven Data Engineer with 8+ years of experience designing, optimizing, and automating data pipelines and data lakes in large-scale enterprise environments. I excel at building scalable batch and streaming data solutions using Snowflake, AWS, Python, dbt, PySpark, Apache Iceberg, Airflow, and AWS Glue to enable real-time analytics, BI, and reliable cloud data ecosystems. I have strong experience in data quality frameworks, governance, and performance tuning across AWS Redshift and Delta Lake, enabling observability, cost efficiency, and reliability while collaborating with analytics, BI, and ML teams to deliver trusted datasets for forecasting and decision-making.

Available to hire

I am a results-driven Data Engineer with 8+ years of experience designing, optimizing, and automating data pipelines and data lakes in large-scale enterprise environments. I excel at building scalable batch and streaming data solutions using Snowflake, AWS, Python, dbt, PySpark, Apache Iceberg, Airflow, and AWS Glue to enable real-time analytics, BI, and reliable cloud data ecosystems.

I have strong experience in data quality frameworks, governance, and performance tuning across AWS Redshift and Delta Lake, enabling observability, cost efficiency, and reliability while collaborating with analytics, BI, and ML teams to deliver trusted datasets for forecasting and decision-making.

See more

Experience Level

Expert
Expert
Expert
Expert
Expert
Expert
Expert
Expert

Work Experience

Sr. Data Engineer at PCI Energy Solutions
October 1, 2024 - Present
Led the design and implementation of end-to-end data ingestion and processing pipelines on AWS using PySpark on EMR and Apache Iceberg, automating extraction, transformation, and loading of large Parquet datasets from multiple sources into S3-backed Iceberg tables to enable scalable analytics. Built batch and near real-time ingestion pipelines with Spark DataFrames, AWS Kinesis, and Structured Streaming, reducing ingestion latency by approximately 40%. Designed a Medallion Architecture data lake (Bronze–Silver–Gold) leveraging Databricks, EMR, Iceberg, and Snowflake to standardize raw ingestion, cleansing, and consumption layers for analytics and ML workloads. Developed high-performance ETL pipelines using PySpark, Delta Lake, Auto Loader, and EMR Spark, applying partitioning, bucketing, clustering, and file compaction to improve efficiency and reduce runtime by 35–40%. Orchestrated complex data workflows with Apache Airflow and AWS Step Functions, implementing retries, SLAs, bac
Sr. Data Engineer at Netcracker Technology Solutions
November 1, 2021 - May 1, 2024
Engineered data integration pipelines across Netcracker BSS modules (CRM, Order Management, Billing, Revenue Management) using Snowflake, Snowpark, DBT, PySpark, and AWS Glue to synchronize customer, subscription, and usage data across systems. Migrated and optimized legacy PL/SQL ETL jobs into Snowflake and Spark-based pipelines, improving throughput and scalability for high-volume telecom data. Built modular DBT transformations to standardize BSS/OSS data for reporting and revenue assurance. Implemented AWS orchestration using Airflow, Lambda, and Glue to automate batch and streaming ETL workflows, including event-driven processing of activations and billing events. Created ETL frameworks to extract customer, subscription, and usage data from BSS/OSS databases into analytic datasets for revenue assurance, activation tracking, and churn analysis. Configured AWS Glue Crawlers and Data Catalog for schema inference and metadata management; provided production support with root cause anal
Data Engineer at Netcracker Technology Solutions
January 1, 2020 - November 1, 2021
Engineered PySpark pipelines to process multi-terabyte CDR and usage data, improving processing speed, scalability, and data reliability. Automated ingestion of BSS/OSS data from relational databases and APIs into AWS S3 and Redshift, enabling near real-time reporting and analytics for telecom operations. Tuned complex SQL queries with window functions and indexing, achieving over 40% faster query performance for billing and customer analytics. Integrated structured and semi-structured data sources, performing data cleansing, transformation, and validation to ensure data quality for BI and ML analytics. Implemented real-time data processing and streaming analytics for telecom networks using Kafka and Spark Streaming, optimizing consumer performance and reducing event processing latency. Collaborated with a team of developers to build a churn prediction data pipeline using PySpark, delivering Power BI dashboards for executive decision-making. Delivered production-ready pipelines with ve
Associate Software Engineer at Netcracker Technology Solutions
August 1, 2017 - December 1, 2019
Designed and implemented relational database schemas aligned with client-specific data models and business requirements to support scalable application design. Developed and executed DDL scripts for table creation, indexing, partitioning, and schema migration across development, testing, and production environments, ensuring seamless and error-free deployment cycles. Created detailed ER diagrams, data dictionaries, and enforced referential integrity constraints to maintain data accuracy and consistency across distributed systems. Defined primary keys, foreign keys, and check constraints to optimize database integrity and enforcement of validation rules. Automated reporting, data extraction, and system maintenance tasks using Shell scripting and Cron, achieving 98% process reliability and reduced manual intervention. Engineered and maintained CI/CD pipelines using Jenkins and GitHub Actions to automate build, test, and deployment workflows across cloud-based environments (AWS EC2, S3).

Education

Bachelors in Computer Science and Engineering at Jawaharlal Nehru Technological University, Hyderabad, India
August 1, 2014 - May 1, 2018

Qualifications

Add your qualifications or awards here.

Industry Experience

Software & Internet, Telecommunications, Professional Services