I'm Farah Gouda, a data engineer specializing in NLP and language model development with a track record of building scalable data pipelines and enabling cross-functional collaboration. I enjoy turning complex data problems into practical solutions, driving AI commercialization, and delivering measurable ROI through language and model improvements.

Farah Gouda

I'm Farah Gouda, a data engineer specializing in NLP and language model development with a track record of building scalable data pipelines and enabling cross-functional collaboration. I enjoy turning complex data problems into practical solutions, driving AI commercialization, and delivering measurable ROI through language and model improvements.

Available to hire

I’m Farah Gouda, a data engineer specializing in NLP and language model development with a track record of building scalable data pipelines and enabling cross-functional collaboration.

I enjoy turning complex data problems into practical solutions, driving AI commercialization, and delivering measurable ROI through language and model improvements.

See more

Experience Level

Expert
Expert
Expert
Expert
Expert
Expert
Expert
Intermediate
Intermediate
Intermediate
See more

Language

English
Fluent
Urdu
Advanced
Bengali
Advanced
Swahili
Advanced

Work Experience

Data Engineer at Speechmatics
August 1, 2024 - November 2, 2025
Integrated 3 new languages (Urdu, Bengali, Swahili) into the ASR pipeline; developed and deployed a Welcome Voice agent for Flow showcased at CES; built an LLM testing pipeline; created interactive dashboards with Plotly Dash for nontechnical stakeholders; designed automated storage reduction scripts for stale experiments; debugged Bash and Python pipelines to resolve CUDA memory bottlenecks, optimize GPU utilization, and improve ASR decoding accuracy; evaluated storage performance (NFS vs S3) and optimized pipelines; developed automated benchmarking to compare ASR accuracy of competitor models on medical datasets across multiple languages; led cross-functional collaboration with PMs and sponsors to drive ASR commercialization and prioritize languages by ROI.
Data Engineer / Developer at Office for National Statistics
February 1, 2024 - February 1, 2024
Lead developer: Built and led the development of two successful pipelines in PySpark processing UK’s Trade; improved published figures by reflecting additions and removals of unrecorded traded goods; transitioned pipelines from Cloudera (Development) to Data Access Platform (DAP) cloud environment; developed ETL pipeline to transfer and transform data to centralized storage using Hive (Data Warehousing) and SQL; maintained best practices through comprehensive testing, RAP principles, documentation and more; regularly presented technical information to non-technical end users for project milestones (discovery, alpha, and beta gates); translated user requirements into technical specifications; demonstrated leadership by line managing and upskilling two junior developers.
Data Scientist at Synapse Analytics
September 1, 2021 - September 1, 2021
Led data analytics for a major e-payments provider in Egypt, creating impactful visualizations and statistical analyses with Matplotlib and Seaborn; identified the root cause of POS-failed transactions, leading to significant annual revenue loss mitigation; conducted thorough EDA, ML model comparison, and network visualization using Pyvis to identify potential fraud and colluding users; effectively communicated data insights to non-technical stakeholders using storytelling techniques; developed AI-based solutions to address patterns and prevent future revenue loss, demonstrating strong business acumen.

Education

MSc Data Science at University of Sussex
October 1, 2020 - July 1, 2021

Qualifications

Add your qualifications or awards here.

Industry Experience

Software & Internet, Media & Entertainment, Government, Professional Services