secure on-premise Hybrid RAG system
We are looking for an expert in AI/ML and web development to build a secure on-premise Hybrid RAG (Retrieval-Augmented Generation) system for a personal project. The system will utilize Ollama or vLLM to run local open-source language models (LLMs), supporting models like Qwen, Llama, Mistral, or DeepSeek for chat, and a multilingual embedding model for Dari, Pashto, and English. You should implement a RAG framework such as LlamaIndex or LangChain, support backend operations in Python FastAPI, and design the frontend using React, Next.js, or Open WebUI. The project requires robust security, including MFA, RBAC, document-level permissions, audit logs, encryption, and backup, as well as enterprise-level deployment with Docker, Ubuntu Server, and future Kubernetes setup. Data will be stored and managed on PostgreSQL and MinIO or other secure storage options, with authentication integration via Keycloak or Active Directory.
The main requirement is for the AI to only answer using approved organizational documents, clearly display the source or page, and strictly enforce user access permissions. The project can be completed fully remotely. If you have experience in DevOps, document processing, and building secure, scalable AI solutions using these technologies, I would like to work with you.
Budget range:
$3,000 โ 5,000 USD
Examples:
a secure on-premise Hybrid RAG system with these technologies:
LLM Runtime: Ollama or vLLM for running local open-source models.
AI Models: Qwen / Llama / Mistral / DeepSeek for chat; multilingual embedding model for Dari, Pashto, English.
RAG Framework: LlamaIndex or LangChain.
Vector Database: Qdrant or PostgreSQL + pgvector.
Backend: Python FastAPI.
Frontend: React / Next.js or Open WebUI for quick start.
Database: PostgreSQL for users, documents, logs, permissions.
Authentication: Keycloak or Active Directory integration with role-based access.
Storage: MinIO or secure internal file storage.
Security: MFA, RBAC, document-level permission, audit logs, encryption, backup.
Deployment: Ubuntu Server, Docker, later Kubernetes for enterprise production.
The key requirement is: AI must answer only from approved Organization documents, show source/page, and respect user access permissions.
- Hiring as
- Individual
- Hiring stage
- Planning and researching
- Work type
- Single job, potential follow-up
- Experience level
- Mid-level
No longer accepting applications
Get instant notifications for new AI Engineer jobs. Enter your email:
How It Works
๐Get quality leads
Review job leads for free, filter by local or global clients, and get real time notifications for new opportunities.
๐Apply with ease
Pick the best leads, unlock contact details, and apply effortlessly with Twine's AI application tools.
๐Grow your career
Showcase your work, pitch to the best leads, land new clients and use Twineโs tools to find more opportunities.
Ari: your job agent
Ari searches hundreds of sources for you, every hour, including company careers pages: more than 10 million postings a month.