DocAssist
A local document Q&A app. Upload PDFs or text files and ask questions about them. Answers are generated by Llama 3.1 8B running on your GPU and include the source document.
Features
- Drag-and-drop upload for PDF, TXT, and MD files
- Answers with source citations
- Runs fully offline with Ollama, no API keys needed
- Five search backends you can switch between: Python, NumPy, C++, multithreaded C++ (OpenMP), and GPU (PyTorch CUDA)
- Benchmarks for search speed and LLM speed
Tech Stack
- Frontend: Streamlit
- LLM: Llama 3.1 8B via Ollama
- Embeddings: nomic-embed-text
- Search: NumPy, C++17 with pybind11 and OpenMP, PyTorch CUDA
- Other: pypdf
How It Works
- Uploaded files are split into overlapping chunks (800 characters, 150 overlap)
- Each chunk is converted to a 768-dimensional embedding
- Questions are embedded the same way and compared to every chunk with cosine similarity
- The top 4 chunks are sent to the LLM, which answers using only that context
Installation
Requirements
- Windows 10/11
- Python 3.12+
- NVIDIA GPU
- Ollama
- Visual Studio Build Tools (Desktop development with C++)
Setup
# Clone the repo
git clone https://github.com/oracl-ee/docassist.git
cd docassist
# Download the models
ollama pull llama3.1:8b
ollama pull nomic-embed-text
# Create a virtual environment
python -m venv .venv
.venv\Scripts\activate
# Install dependencies
python -m pip install -r requirements.txt
python -m pip install torch --index-url https://download.pytorch.org/whl/cu128
# Build the C++ extension
python setup.py build_ext --inplace
Usage
Start the web app:
python -m streamlit run app_ui.py
Command-line version:
python main.py ask "What is a warp?" --backend gpu
python main.py chat
Run tests and benchmarks:
python tests/test_backends.py
python -m bench.bench_search
python -m bench.bench_llm
LLM speed
| Model | Tokens/sec |
|---|---|
| llama3.1:8b (GPU) | 88.5 |
| llama3.1:8b (CPU) | TBD |
| llama3.2:3b (GPU) | TBD |
Project Structure
docassist/
├── app/
│ ├── loader.py # Reads PDF, TXT, and MD files
│ ├── chunker.py # Splits text into chunks
│ ├── embedder.py # Creates embeddings
│ ├── store.py # Saves and loads the index
│ ├── search.py # Search backends
│ └── llm.py # LLM prompts and streaming
├── cpp/
│ └── fastsearch.cpp # C++ search module
├── bench/ # Benchmarks
├── tests/ # Tests
├── app_ui.py # Streamlit app
├── main.py # CLI
└── setup.py # Builds the C++ extension
Troubleshooting
- "An Application Control policy has blocked this file": use
python -m pipandpython -m streamlitinstead ofpipandstreamlit. - First answer is slow: the model is loading into VRAM. After that, it stays loaded for 30 minutes.
Roadmap
- Approximate nearest neighbor search (HNSW / FAISS) for large document sets
- Better chunking (split on sentences instead of characters)
- Evaluation set to measure answer accuracy
License
MIT