I’m an AI/ML and Backend engineer focused on building production-grade GenAI systems—especially RAG pipelines, multi-agent architectures, and low-latency LLM inference. I’ve worked across the stack to ship real features for users, with a strong emphasis on measurable performance (latency, throughput, and cost) and reliable experimentation. I’m comfortable taking models from research to deployment, including quantization and inference optimization, and then wrapping them with robust APIs, telemetry, and dashboards. I enjoy engineering systems that go beyond demos—routing requests intelligently, tracking quality and drift, and continuously improving evaluation results to deliver dependable AI-powered products.

Vedant Kene

I’m an AI/ML and Backend engineer focused on building production-grade GenAI systems—especially RAG pipelines, multi-agent architectures, and low-latency LLM inference. I’ve worked across the stack to ship real features for users, with a strong emphasis on measurable performance (latency, throughput, and cost) and reliable experimentation. I’m comfortable taking models from research to deployment, including quantization and inference optimization, and then wrapping them with robust APIs, telemetry, and dashboards. I enjoy engineering systems that go beyond demos—routing requests intelligently, tracking quality and drift, and continuously improving evaluation results to deliver dependable AI-powered products.

Available to hire
See more

Experience Level

Expert
Expert
Expert
Expert
Expert
Expert
Expert
Expert
Expert
Expert
Intermediate
Beginner
Beginner
See more

Language

English
Advanced
Hindi
Advanced

Work Experience

Deep Learning Inference Intern at IIT Kanpur
July 1, 2025 - December 1, 2025
Applied INT4/INT8 post-training quantization to vision transformer models using PyTorch and NVIDIA TensorRT, achieving ~4x memory reduction with less than 1% accuracy degradation. Built GPU-aware compression pipelines automating the full quantization workflow (model loading, calibration, engine export, and benchmarking), reducing turnaround time from hours to minutes. Designed benchmarking scripts to evaluate latency/throughput/accuracy tradeoffs and tracked experiments using structured Git branching for reproducible research.
ML & Data Engineering Intern at Edvantange Pvt. Ltd., New Delhi
September 1, 2024 - April 1, 2025
Built Python-based RAG pipelines and REST API endpoints serving an AI-powered grading tool for a React frontend, reducing manual assessment effort by 60%. Automated data pipelines using AWS Lambda and REST APIs, cutting manual reporting effort by 45%, and delivered Power BI dashboards for management decision-making. Designed MongoDB collections and schemas for student data, model outputs, and session history. Collaborated in an Agile sprint workflow using Git feature-branch strategy, with structured testing and versioned iterations.
ML & Data Engineering Intern at Eduvantage Pvt. Ltd.
September 1, 2024 - April 1, 2025
Built Python-based RAG pipelines and REST API endpoints powering an AI grading assistant used with a React frontend, reducing manual assessment effort by ~60%. Developed automated data pipelines using AWS Lambda and REST APIs, cutting manual reporting effort by ~45% and delivering Power BI dashboards for management decision-making. Designed MongoDB collections and SQL schemas to store student data, model outputs, and session history, collaborating in an Agile sprint workflow using Git feature-branching strategy. Also contributed to deep learning inference support and engineering practices for reliable experimentation.

Education

B.Tech (Electronics & Telecommunication) at Vishwakaramha Institute of Information Technology
January 1, 2022 - January 1, 2026
B.Tech, Electronics & Telecommunications at Vishwakaram Institute of Information Technology, Pune
January 1, 2022 - January 1, 2026

Qualifications

Add your qualifications or awards here.

Industry Experience

Education, Software & Internet, Telecommunications, Computers & Electronics