Staff AI/ML Engineer with 12 years of experience designing, deploying, and operating production machine learning systems. I build low-latency inference services with robust model monitoring, versioning, and performance regression testing, and I scale core AI products with cloud infrastructure on AWS and GCP. I also deliver enterprise GenAI platforms and agentic systems, including secure RAG architectures and AI governance frameworks for production AI workloads. My work spans PyTorch/TensorFlow, Docker/Kubernetes, high-throughput data pipelines, SQL-driven backend integration, and reliability-focused MLOps practices.

Hassan Kianinejad

Staff AI/ML Engineer with 12 years of experience designing, deploying, and operating production machine learning systems. I build low-latency inference services with robust model monitoring, versioning, and performance regression testing, and I scale core AI products with cloud infrastructure on AWS and GCP. I also deliver enterprise GenAI platforms and agentic systems, including secure RAG architectures and AI governance frameworks for production AI workloads. My work spans PyTorch/TensorFlow, Docker/Kubernetes, high-throughput data pipelines, SQL-driven backend integration, and reliability-focused MLOps practices.

Available to hire

Staff AI/ML Engineer with 12 years of experience designing, deploying, and operating production machine learning systems. I build low-latency inference services with robust model monitoring, versioning, and performance regression testing, and I scale core AI products with cloud infrastructure on AWS and GCP.

I also deliver enterprise GenAI platforms and agentic systems, including secure RAG architectures and AI governance frameworks for production AI workloads. My work spans PyTorch/TensorFlow, Docker/Kubernetes, high-throughput data pipelines, SQL-driven backend integration, and reliability-focused MLOps practices.

See more

Experience Level

Work Experience

Staff Software Engineer at Twelve Labs
November 1, 2023 - Present
Own production AI/ML platform for multimodal deep learning services. Deploy scalable and reliable inference APIs on AWS and GCP using PyTorch/TensorFlow/Python, Docker, and Kubernetes with automated rollouts, health checks, observability dashboards, and versioned model artifacts. Design secure RAG architectures for enterprise GenAI, enabling agentic systems to retrieve from proprietary knowledge bases while enforcing governance, access controls, audit trails, tenant isolation, and enterprise security/compliance requirements. Lead LLMOps adoption via model versioning, automated retraining pipelines, and monitoring dashboards tracking drift, latency, and reliability using Python and SQL. Build high-throughput data pipelines in Python/SQL for multimodal training/evaluation datasets integrated with cloud object storage, feature stores, and backend APIs for continuous improvement. Optimize transformer inference with quantization, dynamic batching, and hardware-aware optimizations to reduce
Senior Deep Learning Software Engineer at Intel Corporation
January 1, 2022 - November 1, 2023
Developed and optimized deep learning training and inference systems for Intel AI accelerators, enabling efficient PyTorch and TensorFlow execution across cloud and on-prem production environments using performance profiling and framework-level enhancements. Implemented quantization and graph optimization to reduce inference latency and improve throughput for computer vision and NLP workloads deployed at scale on AWS and GCP, with automated regression testing and accuracy validation. Built MLOps pipelines with Docker and Kubernetes for automated model testing, versioning, and deployment across cloud platforms integrated with SQL-based model metadata stores. Designed Python backend services and REST APIs to expose optimized inference engines, integrating with SQL databases for model metadata/configuration/performance metrics and supporting real-time traffic. Collaborated with researchers/engineers to profile bottlenecks, conduct code reviews, and deliver hardware-software co-design impr
Software Engineer at DeepRoute.ai
January 1, 2021 - January 1, 2022
Engineered deep learning inference systems for autonomous driving, deploying PyTorch models to cloud and embedded infrastructure with optimized low-latency pipelines and service layers for real-time decision making. Developed data processing and annotation pipelines in Python/SQL to manage large-scale sensor datasets supporting training/validation of perception models with consistent quality control and efficient storage on AWS and GCP. Built internal MLOps tooling with Docker/Kubernetes to standardize deployments, monitor runtime performance, and enable seamless rollback across multiple vehicle platforms to reduce downtime. Implemented REST APIs and Python backend services integrating perception outputs with vehicle control systems under strict latency requirements, using SQL-based configuration management and runtime monitoring. Optimized inference throughput using quantization, layer fusion, and memory management for resource-constrained autonomous vehicle hardware while maintaining
Senior Deep Learning Software Engineer at NVIDIA
December 1, 2017 - January 1, 2021
Developed and optimized deep learning frameworks and inference systems for NVIDIA accelerated computing platforms, improving execution efficiency of PyTorch and TensorFlow workloads via systematic performance analysis and optimization. Built production-grade MLOps pipelines with Docker/Kubernetes and cloud infrastructure on AWS/GCP to deploy models at scale with continuous monitoring, versioning, and rollback. Designed Python tools and APIs for benchmarking model performance, latency, and throughput on hardware generations; used SQL-backed storage and reporting to support deployment recommendations and capacity planning. Implemented model versioning and experiment tracking with SQL backends to enable reproducible research, reliable production handoffs, and auditable trails across teams/products. Accelerated CV/NLP models using mixed precision training, distributed execution, and operator fusion to reduce time-to-deploy and improve inference performance for customers. Collaborated with
Staff Software Engineer at Baidu USA
December 1, 2015 - December 1, 2017
Led development of large-scale ML training and serving infrastructure using Python/SQL/distributed systems for deep learning models supporting search, recommendation, and NLP with high reliability and scalability. Created scalable data pipelines in Python/SQL to ingest, clean, and transform massive datasets enabling frequent retraining, feature engineering, and real-time updates across production services on AWS/GCP. Designed and deployed RESTful model inference APIs and backend services integrating cloud infrastructure and SQL databases to serve millions of low-latency/high-availability requests. Implemented monitoring and operational dashboards to track reliability, latency, and prediction quality with automated alerting and incident response. Championed code review/testing/documentation best practices across cross-functional teams for maintainability and long-term reliability. Optimized serving infrastructure through load balancing, caching, and asynchronous processing to reduce res

Education

M.S., Electrical Engineering at UCLA
January 1, 2013 - January 1, 2015
B.S., Electrical Engineering at University of Tehran
January 1, 2009 - January 1, 2013

Qualifications

Add your qualifications or awards here.

Industry Experience

Software & Internet, Computers & Electronics, Professional Services