I’m a Senior AI Engineer who builds production-grade LLM-powered systems, specializing in RAG (retrieval-augmented generation) and agentic AI architectures. I’ve led end-to-end development across data pipelines, model training, retrieval and orchestration, evaluation, and observability—designing solutions that are reliable, low-latency, and cost-efficient for real-world business needs. From legal summarization engines to enterprise conversational AI and fraud/risk modeling, I translate complex, ambiguous domain problems into robust, measurable products. I work across AWS and GCP, using modern ML/MLOps and LLMOps tooling to deploy scalable inference, validate output quality, and continuously improve performance through automated benchmarks and monitoring.

Mihai Dragomir

I’m a Senior AI Engineer who builds production-grade LLM-powered systems, specializing in RAG (retrieval-augmented generation) and agentic AI architectures. I’ve led end-to-end development across data pipelines, model training, retrieval and orchestration, evaluation, and observability—designing solutions that are reliable, low-latency, and cost-efficient for real-world business needs. From legal summarization engines to enterprise conversational AI and fraud/risk modeling, I translate complex, ambiguous domain problems into robust, measurable products. I work across AWS and GCP, using modern ML/MLOps and LLMOps tooling to deploy scalable inference, validate output quality, and continuously improve performance through automated benchmarks and monitoring.

Available to hire

I’m a Senior AI Engineer who builds production-grade LLM-powered systems, specializing in RAG (retrieval-augmented generation) and agentic AI architectures. I’ve led end-to-end development across data pipelines, model training, retrieval and orchestration, evaluation, and observability—designing solutions that are reliable, low-latency, and cost-efficient for real-world business needs.

From legal summarization engines to enterprise conversational AI and fraud/risk modeling, I translate complex, ambiguous domain problems into robust, measurable products. I work across AWS and GCP, using modern ML/MLOps and LLMOps tooling to deploy scalable inference, validate output quality, and continuously improve performance through automated benchmarks and monitoring.

See more

Language

Work Experience

Senior AI Engineer | Tech Team Lead at LexisNexis (Lexis+ AI)
February 1, 2024 - July 31, 2026
Led development of a production-grade Legal Summarization Engine processing long-form case law (100+ pages) into structured headnotes and holdings, reducing attorney review time by 60%+. Fine-tuned LLaMA 3 70B and LLaMA 2 13B with QLoRA to generate citation-grounded summaries, improving factual consistency by ~45%. Built large-scale SFT pipelines using Apache Spark + Delta Lake to curate and preprocess 10M+ legal documents. Designed hierarchical summarization for documents exceeding 200K tokens. Architected multi-agent workflows with LangGraph (Retriever, Summarizer, Critic) and implemented Model Context Protocol (MCP) for standardized context exchange across agents and retrieval. Integrated GPT-4o and Claude for synthetic data generation, distillation, and LLM-as-judge evaluation. Implemented hybrid RAG with Pinecone + Elasticsearch and a citation validation layer to reach >92% citation accuracy, with automated evaluation and performance optimizations using PyTorch FSDP and vLLM/Tenso
Senior AI Engineer | Tech Team Lead at LexisNexis
February 1, 2024 - July 31, 2026
Led development of a production-grade Legal Summarization Engine on Lexis+ AI, processing long-form case law into structured headnotes and holdings and reducing attorney review time by 60%+. Fine-tuned LLaMA 3 70B and LLaMA 2 13B using QLoRA to generate citation-grounded legal summaries with improved factual consistency. Built large-scale SFT data pipelines using Apache Spark + Delta Lake to curate and preprocess 10M+ legal documents. Designed hierarchical summarization for documents exceeding 200K tokens and implemented multi-agent workflows with LangGraph (Retriever, Summarizer, Critic), reducing post-editing cycles by ~40%. Implemented RAG with hybrid retrieval (Pinecone + Elasticsearch), built citation validation for jurisdiction-specific references (>92% accuracy), and established automated evaluation pipelines for faithfulness and reasoning consistency. Optimized training with PyTorch FSDP and deployment with vLLM and TensorRT-LLM to achieve 2–3x throughput gains and sub-second
Senior AI Engineer at Parloa
December 1, 2021 - January 31, 2024
Led development of enterprise conversational AI systems across voice and chat channels, improving intent recognition accuracy by 30%+ and scaling automated support for high-volume contact center operations. Designed end-to-end real-time pipelines (ASR → NLU → Dialogue Manager → TTS) to reduce average handling time by 20–25% and improve CSAT/containment. Fine-tuned BERT/RoBERTa models for intent classification and NER, improving precision/recall by ~25% across multilingual interactions. Built retrieval-based response systems using Elasticsearch + dense embeddings, and implemented multi-turn dialogue state management and memory to reduce user repetition. Integrated GPT-3.5/early GPT-4 for dynamic responses and guardrails to ensure safe, consistent, brand-aligned outputs, reducing escalation to human agents by ~35%. Optimized low-latency inference with FastAPI and deployed scalable AI services on AWS with CI/CD pipelines.
Senior Data Scientist at Revolut
September 1, 2018 - October 31, 2021
Built and deployed machine learning models for fraud detection and transaction anomaly detection, reducing fraudulent transactions by 30%+ across card payment flows. Developed real-time risk scoring for payment authorization with improved fraud precision while maintaining low false-positive rates. Designed feature engineering pipelines over large-scale financial datasets (billions of transactions), improving performance by 20%+. Implemented gradient boosting models (XGBoost/LightGBM) for credit risk and user behavior prediction, improving approval accuracy and reducing losses. Built real-time streaming analytics pipelines (Kafka + Spark Streaming) for fraud monitoring with sub-second decision latency. Developed experimentation and A/B testing frameworks, customer segmentation using clustering, and ensemble modeling/feature store architecture that increased detection recall by 25%+. Deployed models as scalable Python microservices via REST APIs.
Software Engineer at Fortech
October 1, 2010 - May 31, 2016
Developed and maintained enterprise-grade backend systems and web applications for European clients, improving reliability and reducing production incidents by ~30%. Built scalable RESTful APIs and microservice-style architectures using the Java/Spring ecosystem for high-traffic enterprise workflows. Designed and optimized relational database schemas and SQL queries, reducing query latency by 40%+. Implemented backend services for data processing and workflow automation, improving operational efficiency. Contributed to distributed system components focused on fault tolerance, logging, and observability, and integrated third-party APIs for secure data exchange. Participated in Agile development lifecycle and supported CI/CD adoption for continuous delivery.

Education

M.S in Computer Science at Stanford University
August 1, 2016 - June 30, 2018
B.S in Computer Science at Technical University of Cluj Napoca
August 1, 2006 - July 31, 2010
M.S in Computer Science at Stanford University
August 1, 2016 - June 30, 2018
B.S in Computer Science at Technical University of Cluj Napoca
August 1, 2006 - July 31, 2010

Qualifications

Add your qualifications or awards here.

Industry Experience

Financial Services, Professional Services, Software & Internet, Media & Entertainment, Government, Healthcare