I'm Kiran Kumar, a Software Architect Engineer with 5+ years of experience in designing secure, high-throughput AI systems and distributed backends for enterprise platforms. I specialize in productionizing Generative AI, strict multi-tenant RAG ecosystems, and autonomous Agentic AI workflows. I combine deep expertise in LLM inference optimization (vLLM, TensorRT) and CI/CD MLOps with building event-driven microservices (Kafka, FastAPI). At work I resolve scaling bottlenecks, enforce data isolation, and deploy low-latency AI solutions across AWS EKS. I enjoy turning experiments into production-ready services, from model packaging with LoRA and MLflow to automated CI/CD pipelines and observability dashboards that track latency and throughput for thousands of tenants.

Kiran Kumar

I'm Kiran Kumar, a Software Architect Engineer with 5+ years of experience in designing secure, high-throughput AI systems and distributed backends for enterprise platforms. I specialize in productionizing Generative AI, strict multi-tenant RAG ecosystems, and autonomous Agentic AI workflows. I combine deep expertise in LLM inference optimization (vLLM, TensorRT) and CI/CD MLOps with building event-driven microservices (Kafka, FastAPI). At work I resolve scaling bottlenecks, enforce data isolation, and deploy low-latency AI solutions across AWS EKS. I enjoy turning experiments into production-ready services, from model packaging with LoRA and MLflow to automated CI/CD pipelines and observability dashboards that track latency and throughput for thousands of tenants.

Available to hire

I’m Kiran Kumar, a Software Architect Engineer with 5+ years of experience in designing secure, high-throughput AI systems and distributed backends for enterprise platforms. I specialize in productionizing Generative AI, strict multi-tenant RAG ecosystems, and autonomous Agentic AI workflows. I combine deep expertise in LLM inference optimization (vLLM, TensorRT) and CI/CD MLOps with building event-driven microservices (Kafka, FastAPI).

At work I resolve scaling bottlenecks, enforce data isolation, and deploy low-latency AI solutions across AWS EKS. I enjoy turning experiments into production-ready services, from model packaging with LoRA and MLflow to automated CI/CD pipelines and observability dashboards that track latency and throughput for thousands of tenants.

See more

Experience Level

Expert
Expert
Expert
Expert
Expert
Expert
Expert
Expert

Work Experience

Software Architect Engineer (AI Engineering & MLOps) at Salesforce, Inc.
September 1, 2023 - Present
Resolved HTTP timeout failures during bulk CRM data syncs by decoupling Einstein AI inference into a FastAPI asynchronous queue, reliably routing 12,000+ LLM payloads per minute. Mitigated OOM pod crashes and reduced AWS GPU costs by quantizing internal XGen models from FP16 to INT8 via TensorRT and vLLM, sustaining sub-200ms latency for Service Cloud. Enforced strict multi-tenant data isolation within the Einstein Trust Layer RAG backend by implementing RBAC metadata filtering in Milvus, allowing secure LLM grounding against Data Cloud records. Transitioned experimental Einstein Copilot agents into production Python microservices, exposing a unified API for LangChain-driven prompt routing and autonomous SOQL query execution against client orgs. Managed Hyperforce AWS EKS clusters for NVIDIA Triton inference, tuning custom Kubernetes HPA metrics (based on Kafka lag and GPU utilization) to auto-scale pods across 3 availability zones. Replaced manual model handoffs by building a GitHub A
Software Architect at SAP
March 1, 2020 - July 1, 2022
Developed predictive deep learning models using TensorFlow and AWS SageMaker, processing 5 million historical procurement records to forecast monthly supply chain inventory bottlenecks. Engineered an NLP text-classification engine utilizing Hugging Face Transformers (BERT), automating vendor categorization and reducing manual data-entry errors by 28% within the SAP Ariba module. Established a centralized MLOps tracking architecture via MLflow, providing a unified dashboard for hyperparameter logging and model drift detection. Architected secure RESTful APIs leveraging FastAPI and PostgreSQL to expose machine learning outputs for real-time enterprise dashboards. Refactored legacy feature-extraction algorithms, reducing processing times from 4 hours to under 45 minutes.
Full stack Engineer at Tiger Analytics
May 1, 2019 - March 1, 2020
Programmed scalable ETL data pipelines utilizing Apache Spark and Python, extracting and transforming raw transactional data from 500+ daily retail logs to feed downstream machine learning models. Developed core Python and PostgreSQL backend infrastructure for an internal predictive analytics tool, bridging foundational statistical outputs with a dynamic React.js interface. Integrated a Redis caching layer into the primary API gateway, offloading repetitive queries and reducing dashboard load latency to under 600 ms. Implemented automated backend validation tests using PyTest within a Jenkins CI/CD pipeline, enhancing training-data quality by 35%. Participated in Agile standups and maintained Swagger API documentation for frontend teams.

Education

Master of Science in Computer Science at Auburn University at Montgomery, Alabama, USA
January 11, 2030 - May 1, 2024
Bachelor of Engineering in Electronics and Communication Engineering at Jawaharlal Nehru Institute of Technology, India
January 11, 2030 - September 1, 2020

Qualifications

Add your qualifications or awards here.

Industry Experience

Computers & Electronics, Software & Internet, Professional Services

Experience Level

Expert
Expert
Expert
Expert
Expert
Expert
Expert
Expert

Hire a Full Stack Developer

We have the best full stack developer experts on Twine. Hire a full stack developer today.