I’m Marium Faheem, a Senior Data Engineer and GenAI / Agentic AI platform lead based in Berlin, where I design and deliver production-grade data platforms and AI systems for enterprise clients. Over the past 6+ years, I’ve architected lakehouse and real-time analytics solutions on AWS, GCP, and Azure, and led teams from prototype to production across highly regulated environments. I also specialize in building GenAI experiences—especially RAG and agentic workflows—turning complex data and retrieval challenges into reliable, measurable outcomes. From creating MCP-enabled servers and AI-native CI/CD code review to shipping autonomous analyst agents and executive intelligence platforms, I focus on scalable engineering, governance, and human-in-the-loop processes that help stakeholders trust and adopt AI.

Marium Faheem

I’m Marium Faheem, a Senior Data Engineer and GenAI / Agentic AI platform lead based in Berlin, where I design and deliver production-grade data platforms and AI systems for enterprise clients. Over the past 6+ years, I’ve architected lakehouse and real-time analytics solutions on AWS, GCP, and Azure, and led teams from prototype to production across highly regulated environments. I also specialize in building GenAI experiences—especially RAG and agentic workflows—turning complex data and retrieval challenges into reliable, measurable outcomes. From creating MCP-enabled servers and AI-native CI/CD code review to shipping autonomous analyst agents and executive intelligence platforms, I focus on scalable engineering, governance, and human-in-the-loop processes that help stakeholders trust and adopt AI.

Available to hire

I’m Marium Faheem, a Senior Data Engineer and GenAI / Agentic AI platform lead based in Berlin, where I design and deliver production-grade data platforms and AI systems for enterprise clients. Over the past 6+ years, I’ve architected lakehouse and real-time analytics solutions on AWS, GCP, and Azure, and led teams from prototype to production across highly regulated environments.

I also specialize in building GenAI experiences—especially RAG and agentic workflows—turning complex data and retrieval challenges into reliable, measurable outcomes. From creating MCP-enabled servers and AI-native CI/CD code review to shipping autonomous analyst agents and executive intelligence platforms, I focus on scalable engineering, governance, and human-in-the-loop processes that help stakeholders trust and adopt AI.

See more

Experience Level

Expert
Expert
Expert
Expert
Expert
Expert
Expert
Expert
Intermediate
Intermediate
Intermediate
Intermediate
Intermediate
Intermediate
See more

Language

English
Fluent
German
Beginner

Work Experience

Senior Data Engineer, GenAI & Data Platform Lead at McKinsey & Company
December 1, 2024 - Present
Led the data engineering and platform workstream delivering a production-grade, trial-level clinical trial simulator MVP in 9 months, designed to scale across therapeutic areas. Architected AWS (EKS, ECR) with GitLab CI/CD and Prefect-orchestrated pipelines, plus Snowflake data products and multi-source API integrations powering clinical analytics and probability-of-success models. Built a production Model Context Protocol (MCP) server with semantic similarity to rank relevant historical trials, and shipped an AI-native code review workflow integrated into CI/CD to enforce standards pre-merge. Co-led a cross-functional team through quarterly value releases, enabling human-in-the-loop generation of up to 1,000 trial design variants.
Senior Data Engineer at McKinsey & Company
May 1, 2023 - November 30, 2024
Delivered an enterprise big data and AI platform for a sovereign wealth fund, leading a team (10 members including vendors) to design and run on-premises real-time and batch pipelines for real-estate and investment analytics. Built scalable pipelines across on-premises, AWS, and GCP using Kimball and Data Vault modeling with Terraform IaC, and managed AWS EKS clusters running Airflow, Airbyte, and Jenkins. Also modernized a national municipality’s data platform, improving data quality and enabling 50+ downstream AI use cases; engineered real-time fraud detection with Kafka/Kinesis/Lambda/DynamoDB and SageMaker with sub-5-second executive dashboards. Led multi-client agentic GenAI delivery integrating LangChain RAG and multi-agent workflows across 30+ enterprise data sources, shipping GenAI stack implementations on AWS OpenSearch and Azure OpenAI.
Data Engineer at McKinsey & Company
May 1, 2023 - November 30, 2024
Led an on-premises big data and AI platform initiative for real-time and batch processing supporting real-estate and investment analytics, serving as solution architect from design through production. Built scalable analytics pipelines across on-premises, AWS, and GCP using Kimball and Data Vault dimensional modeling, using Terraform IaC; managed AWS EKS clusters running Airflow, Airbyte, and Jenkins. Delivered data platform modernization and AI enablement for a national municipality (fraud detection and ETL pipelines) and created executive dashboards targeting sub-5-second latency. Also delivered an enterprise agentic GenAI platform integrating LangChain RAG, vector databases, and multi-agent workflows across 30+ enterprise sources, deploying GenAI stacks on AWS OpenSearch and Azure OpenAI.
Data Engineer I at Bazaar Technologies
January 1, 2023 - April 30, 2023
Designed and scaled a high-performance ML platform covering ingestion, transformation, training, and real-time inference, with low-latency prediction APIs for business-critical applications. Established a governed model registry with full versioning and lineage tracking. Automated MLOps CI/CD pipelines for rapid experimentation, deployment, and monitoring at scale, and scaled real-time order-insight infrastructure serving millions of customers, collaborating with Uber and Apple teams to enable EMR on EKS real-time pipelines.
Senior Data Engineer I at Bazaar Technologies
January 1, 2023 - April 30, 2023
Designed and scaled a high-performance machine learning platform covering ingestion, transformation, training, and real-time inference, delivering low-latency prediction APIs for business-critical applications. Established a model registry with full versioning, lineage tracking, governance, and automated MLOps CI/CD pipelines for experimentation, deployment, and monitoring. Scaled a real-time order-insight infrastructure used by millions of customers, collaborating with Uber and Apple teams to enable a real-time pipeline on EMR on EKS.
Data Engineer II at Bazaar Technologies
August 1, 2020 - January 31, 2023
As the first and only data engineer for two years, built the data function end-to-end from scratch. Designed and delivered an analytical platform on Kubernetes and Hadoop, orchestrating batch and streaming pipelines with Airflow, Spark, and Livy over an S3 / Apache Hudi / Trino lake. Architected lakehouse/data mesh patterns and provisioned infrastructure with Terraform, building an AWS warehouse (Redshift, S3, Glue, DMS, MySQL RDS) with governed dimensional models. Implemented Kafka for event streaming and deployed Spark on EC2 and EMR on EKS with Hudi for ACID-compliant storage. Built MLOps tooling on OpenSearch, automated deployments with AWS CodePipeline, managed Kubernetes clusters with Helm, and delivered Python/PySpark transformations and real-time BI dashboards.
Data Engineer II (promoted from Data Engineer I) at Bazaar Technologies
August 1, 2020 - January 31, 2023
Built the company’s analytical data function from scratch as its first data engineer, designing and delivering end-to-end batch/streaming pipelines on Kubernetes and Hadoop with Airflow, Spark, and Livy over an S3-based data lake using Apache Hudi and Trino. Architected scalable data lake, data mesh, and lakehouse architectures; provisioned infrastructure via Terraform and built the AWS warehouse with Redshift, S3, Glue, DMS, and MySQL RDS. Implemented Kafka for event streaming, deployed Spark on EC2 with Docker, and managed EMR on EKS with Apache Hudi for versioned, ACID-compliant storage. Built an MLOps platform on the OpenSearch stack and contributed to Apache Hudi; developed Python/PySpark transformations and real-time BI dashboards for decision-making.

Education

Bachelor of Engineering (Computer Systems) at NED University of Engineering & Technology
January 1, 2017 - January 1, 2020
Bachelor of Engineering (Computer Systems) at NED University of Engineering & Technology, Karachi, Pakistan
January 1, 2016 - October 1, 2020

Qualifications

Databricks Certified Associate Developer for Apache Spark 3.0
January 11, 2030 - August 28, 2026
AWS Certified Cloud Practitioner
January 11, 2030 - August 28, 2026
IBM Data Engineering Professional
January 11, 2030 - August 28, 2026
IBM Data Science Professional
January 11, 2030 - August 28, 2026
Big Data Specialization
January 11, 2030 - August 28, 2026
Introduction to Generative AI
January 11, 2030 - August 28, 2026
Databricks Certified Associate Developer for Apache Spark 3.0
January 11, 2030 - August 28, 2026
AWS Certified Cloud Practitioner
January 11, 2030 - August 28, 2026
IBM Data Engineering Professional
January 11, 2030 - August 28, 2026
IBM Data Science Professional
January 11, 2030 - August 28, 2026
Big Data Specialization
January 11, 2030 - August 28, 2026
Introduction to Generative AI
January 11, 2030 - August 28, 2026

Industry Experience

Professional Services, Healthcare, Financial Services, Government, Life Sciences, Software & Internet, Real Estate & Construction