Doc- RAG (Document based Retrieval-Augmented Generation) is a powerful tool that enables users to upload any PDF document (books, user manuals, reports) and ask questions about its content. The system leverages multimodal processing to understand both text and visual elements within the document, providing comprehensive answers based on the entire context.
- PDF Processing: Upload any PDF document for analysis
- Multimodal Understanding: Processes both text and images within documents
- Natural Language Querying: Ask questions in plain English about any aspect of the document
- Context-Aware Responses: Receives answers that incorporate information from relevant sections
- Interactive Web Interface: User-friendly Streamlit interface for easy document upload and querying
MRAG uses a "Summarization and Descriptive Embedding" approach where:
- PDFs are preprocessed to extract text, images, and tables using Unstructured.io
- A multimodal LLM (Claude 3.7 Sonnet) to generate detailed descriptions for extracted images
- These texts and descriptions are embedded into a single vector database (weaviate)
- Image embeddings are mapped to a unique ID linked to the original content
- Metadata for text and images are stored in a separate database (Mongo DB) to provide context to user
- During inference, the system:
- Vectorizes the user query
- Performs similarity search to retrieve relevant content
- Fetches the corresponding images/metadata using unique IDs
- Generates comprehensive answers using the retrieved context
- Backend: Python with FastAPI
- Frontend: Streamlit for interactive web interface
- MLLM: Claude 3.7 Sonnet (via AWS Bedrock)
- Vector Database: Weaviate
- Key Libraries:
- Unstructured: For extracting elements from PDFs
- LangChain: For implementing retrieval pipelines with ChromaDB
- Weaviate: For efficient similarity search
- MongoDB: For Document and metadata storage
- One way to run the application as a single docker container. For this clone the repo and create a environment file (.env) containing all the environment variables as shown in .env.example file.
git clone https://github.com/prvnsingh/Doc-RAG.git
#create .env with variable as .env.example
docker compose up --build
- Another way is to run the application by setting up individual components: backend, frontend and vector DB, follow the below commands
# Clone the repository
git clone https://github.com/prvnsingh/Doc-RAG.git
cd Doc-RAG
# Create and activate a virtual environment (optional but recommended)
python -m venv venv
source venv/bin/activate # On Windows: venv\Scripts\activate
# Install dependencies
pip install -r requirements.txt
# Set up environment variables
cp .env.example .env
# Edit .env with your AWS credentials and other configuration
# AWS_REGION=#####
# AWS_ACCESS_KEY_ID=################
# AWS_SECRET_ACCESS_KEY=#################
# OCR_AGENT=pytesseract
# TESSERACT_LANGUAGE=eng
# UNSTRUCTURED_HI_RES_MODEL_NAME=yolox
# MONGO_URI=######### (MongoDB URI)
# no need of WEAVIATE_HOST as we will spin it locally
# Spin the docker image of Weaviate to setup local vectorDB
docker compose -f docker-compose-vector-db.yml up --build
To run the application:
Run the back-end framework
# Start the FastAPI server
uvicorn app.main:app --reload
# Access the API
# Open your browser and go to http://localhost:8000/docsRun the frontend
# Start the Streamlit app
streamlit run app/streamlit_app.py
# Access the web interface
# Open your browser and go to http://localhost:8501src/
├── app/
│ ├── main.py # FastAPI application
│ ├── streamlit_app.py # Streamlit interface
│ ├── config.py # Configuration settings
│ └── prompt.py # Prompt templates
├── components/ # Reusable components
├── services/ # Core services
├── resources/ # Static resources
└── settings.py # Project settings
- Implement an ensembled extraction pipeline using tools like Pix2Text for more accurate extraction of images, tables, and mathematical equations
- Develop a better ranking mechanism using re-ranking models
- Add support for more document formats beyond PDF
- Implement user authentication and document management
- Add batch processing capabilities for multiple documents
Contributions are welcome! Please feel free to submit a Pull Request.
This project is licensed under the MIT License - see the LICENSE file for details.
