Demo video·Watch on Google Drive
NexusRAG-Advanced-Document-Intelligence
An advanced, full-stack Retrieval-Augmented Generation (RAG) application built by Nipun Abhilash. This project allows users to upload complex PDF documents and instantly query them using the state-of-the-art Google Gemini 3.6 Flash reasoning model.
About this project
Building a standard RAG (Retrieval-Augmented Generation) app is easy. Building a resilient, production-grade RAG system that doesn't crash on massive documents? That’s a completely different engineering challenge. I’m thrilled to unveil my latest project: NexusRAG: Advanced Document Intelligence 🧠 (Check out the 1-minute demo !) While building this, I realized relying on standard out-of-the-box API wrappers wasn't cutting it for enterprise-level document processing. So, I ripped out the foundation and built a system that truly pushes the limits of what LLMs can do with complex data. Here is why NexusRAG stands out as a next-level architecture: 🛠️ Zero Data-Loss Custom Embedding Engine: I hit a wall with legacy SDK bugs causing SSL EOF crashes during massive vectorization. My solution? I engineered a custom HTTP-level embedding engine using httpx. It features built-in exponential backoff and rate-limit bypassing, ensuring large PDFs are chunked and vectorized flawlessly without a single dropped connection. 🧠 Deep Reasoning with Gemini 3.6 Flash: I didn't just want a bot that summarizes; I wanted a bot that thinks. By piping LangChain into Google’s Gemini 3.6 Flash and unlocking a massive 8,192 token output limit, NexusRAG utilizes Chain-of-Thought (CoT) reasoning to deliver highly detailed answers that span across dozens of pages. 🔍 Maximum Context Retrieval: Using a dynamic InMemoryVectorStore retriever tuned to k=20, the AI is fed massive amounts of context, ensuring it never loses the plot—no matter how complex the user's query is. ✨ Premium Glassmorphic UI: Functionality shouldn't compromise design. I built a fully responsive, custom-styled glassmorphic interface using Gradio and native CSS to deliver a sleek user experience. ⚙️ The Tech Stack: Python | LangChain | Google Gemini 3.6 Flash | Gradio | Custom HTTP Integrations I am currently benchmarking the vector retrieval latency locally before I containerize the backend with Docker for a full cloud deployment. Turning theoretical machine learning models into highly practical, resilient software tools is what I love doing most.


