Point at a folder, get a clean table of file text + metadata back — 100% offline.
Scrubkit walks a directory and extracts text and metadata from common file types (PDF, Word, Excel, PowerPoint, Email, OpenDocument, EPUB, Plain Text, EXIF), returning one row per file.
Explore complete Scrubkit documentation, integration recipes, and API references on our web portal:
| Resource | Description | Destination Link |
|---|---|---|
| 🌐 Official Website | Interactive product overview, performance benchmarks, and architecture. | Visit Site → |
| 🤖 RAG & AI Recipes | End-to-end recipes for Microsoft.Extensions.AI, Semantic Kernel, vector stores, and Parquet. |
View RAG Recipes → | Markdown |
| 💻 CLI Command Guide | Zero-code folder scanning via scrubkit scan, output formatting, and CI secret scanning. |
View CLI Guide → |
| 📦 Package Family | Full matrix of all 12 Scrubkit packages (Core, Abstractions, Email, Parquet, etc.). |
Explore Packages → |
| 🧪 Interactive Demo | Runnable demo console application on synthetic sample files. | Run Playground → |
| 📝 Release History | Version changelogs, feature additions, and API update notes. | View Changelog → |
- License: Mozilla Public License 2.0 (MPL-2.0) — open and free to use in open or commercial applications.
- Privacy: 100% offline. No network requests or telemetry.
