I’m Jeremy Sigamony, an AI Engineer based in Islamabad, focused on building practical, production-ready generative AI and agentic voice solutions. I enjoy turning complex workflows into reliable systems—whether it’s automating end-to-end booking and ordering over voice, or making RAG and semantic search fast and accurate using strong retrieval pipelines. Across my work, I’ve shipped multi-agent architectures with LangChain/LangGraph, migrated and optimized vector search for hybrid retrieval, and refactored AI generation pipelines into clean microservices with webhook-driven orchestration. I’m passionate about MLOps, backend systems, and conversational AI—always aiming for measurable improvements like lower latency, better response quality, and higher conversion.…

Jeremy Sigamony

I’m Jeremy Sigamony, an AI Engineer based in Islamabad, focused on building practical, production-ready generative AI and agentic voice solutions. I enjoy turning complex workflows into reliable systems—whether it’s automating end-to-end booking and ordering over voice, or making RAG and semantic search fast and accurate using strong retrieval pipelines. Across my work, I’ve shipped multi-agent architectures with LangChain/LangGraph, migrated and optimized vector search for hybrid retrieval, and refactored AI generation pipelines into clean microservices with webhook-driven orchestration. I’m passionate about MLOps, backend systems, and conversational AI—always aiming for measurable improvements like lower latency, better response quality, and higher conversion.…

Available to hire

I’m Jeremy Sigamony, an AI Engineer based in Islamabad, focused on building practical, production-ready generative AI and agentic voice solutions. I enjoy turning complex workflows into reliable systems—whether it’s automating end-to-end booking and ordering over voice, or making RAG and semantic search fast and accurate using strong retrieval pipelines.

Across my work, I’ve shipped multi-agent architectures with LangChain/LangGraph, migrated and optimized vector search for hybrid retrieval, and refactored AI generation pipelines into clean microservices with webhook-driven orchestration. I’m passionate about MLOps, backend systems, and conversational AI—always aiming for measurable improvements like lower latency, better response quality, and higher conversion.

See more

Work Experience

AI Engineer at Infinite Forces
April 1, 2026 - Present
Built voice AI solutions with agentic capabilities to automate end-to-end booking, scheduling, and ordering workflows, reducing manual handling of routine customer requests. Designed conversational agents to answer inbound customer enquiries autonomously and escalate edge cases appropriately. Led integration work connecting voice agents with backend booking/ordering systems and third-party APIs to enable seamless real-time transaction handling. Automated Meta Ads lead response with an AI voice assistant to call every new lead within 60 seconds of submission, boosting conversion rate by 45%.
AI Techfellow at Taleemabad
July 1, 2025 - August 1, 2025
Built and iterated on an NL2SQL chatbot pipeline enabling natural-language querying over structured databases, improving non-technical data retrieval workflows.
AI Engineer at CCRIPT Agency
July 1, 2025 - March 1, 2026
Improved RAG chatbot response latency by 60% and answer quality by overhauling the document chunking strategy and migrating to Elasticsearch as the vector store, enabling Hybrid Search (BM25 + ANN). Integrated WhatsApp delivery via Twilio, decoupled the chatbot into a standalone Python FastAPI microservice from the core Node.js backend, and deployed the full stack on a DigitalOcean instance. Refactored a monolithic NestJS AI generation pipeline into discrete FastAPI microservices (TTS/voice cloning, video generation, and talking-head/slideshow), replacing inefficient polling with webhook-driven callbacks and offloading media persistence to Cloudflare R2 for better scalability. Engineered semantic search for a Shopify storefront using AWS Bedrock embeddings with a dual-layer vector store (ChromaDB for dev, S3-backed for production) targeting sub-300ms query latency, plus cron-based ingestion for re-indexing product listings. Developed a FastAPI backend for a construction estimation plat
AI Intern at CareCloud
April 1, 2025 - June 1, 2025
Architected a multi-agent AI workflow using LangChain and LangGraph integrating web search and database retrieval tools for dynamic, context-aware query resolution. Evaluated and benchmarked open-source STT models (Whisper, Vosk, Moonshine) for real-time transcription accuracy and latency across clinical audio scenarios.

Education

BS, Computer Science at FAST NUCES (National University of Computer and Emerging Sciences)
January 1, 2021 - January 1, 2025

Qualifications

Add your qualifications or awards here.

Industry Experience

Software & Internet, Education, Healthcare, Professional Services, Media & Entertainment, Computers & Electronics